AI application security: model capabilities and system boundaries.
Security depends on data, access, tools and permitted actions. A more capable model calls for reassessing the assumptions used to protect a system.
Available capabilities change
NIST’s September 2026 assessment of GLM-5.3 shows advanced cyber capabilities in an open-weight model, still below the US frontier on aggregate tests. It signals a spread of capabilities, rather than the frequency of real attacks.
The implication I draw is to revisit the threat model periodically: who can attack the system, with which tools, which data or functions they can reach and what damage they can cause.
For an exposed service, I therefore give even more attention to vulnerability management, dependency updates and credential protection. The same AI tools can also help defenders, but their findings need verification.
Permissions must stay in the application
Model instructions are part of the defence; the decisive controls must be enforced by the system.
A document, retrieved page or tool output can contain hostile instructions. Prompt injection exploits this confusion between content and instructions. The workflow must treat these elements as untrusted data.
Authorisation and user separation must be checked in code. A model's proposal must not grant new privileges or provide access to data the user could not consult directly.
Limit agent actions
For every tool connected to the model, I define what it can do and with which credentials. Minimum permissions, bounded operations and separate environments reduce the consequences of mistaken or hostile requests.
- Validate parameters and destinations before performing an action.
- Require confirmation or risk-proportionate review for consequential operations.
- Set limits on attempts, duration and consumption, with a stopping condition.
- Log actions with attention to sensitive data, so incidents can be reconstructed.
These controls belong in the application workflow. An agent that can use more tools should not automatically receive more authority.
Assess the actual configuration
Open-weight models and proprietary services offer different controls and risks. Available weights do not alone make a deployment unsafe, just as a proprietary API does not alone guarantee product security.
I assess the model's origin and version, runtime, dependencies, network, credentials, access and data path. A local environment still needs maintenance, updates and user access controls.
Verify and reassess over time
Alongside functional tests, I use hostile requests, unauthorised access attempts and retrieved content that tries to divert the workflow. I focus on whether boundaries remain effective when the model makes a mistake.
Changes to models, tools, prompts and data require new checks. Monitoring, reporting and incident response should be planned before release. This is an ongoing process, proportionate to the consequences of errors.
Supporting sources
The cyber assessments describe the cited models and tests. OWASP references support the application risks; operational choices must be adapted to the actual system.
- NIST CAISI — GLM-5.3 (opens in a new tab): cyber capabilities and limits of the frontier comparison.
- Anthropic — spread of cyber capabilities (opens in a new tab): exploit and safeguard assessments in test environments.
- OWASP — Prompt Injection (opens in a new tab): risks from hostile instructions in content processed by a model.
- OWASP — Excessive Agency (opens in a new tab): excessive actions, permissions and autonomy in systems with tools.