Resume
An AI assistant connected to your data can disclose confidential information to users who are not authorized to access it. Simply instructing the assistant to keep that information confidential is not enough: as long as it receives the data, it may include it in its responses. Protecting sensitive data starts with clearly defined access controls that are enforced before the data is sent to the model. Instructions should complement these controls, not replace them.
AI assistants and agents based on large language models (LLMs) can access documents, query databases and use internal tools. This makes information easier to access, but it can also create a path to data that some users are not authorized to view.
A vulnerability disclosed in 2025 in Lena, Lenovo’s customer service assistant, illustrates this risk. Cybernews researchers showed that a message of roughly 400 characters could cause the assistant to include malicious code in its response. When a support agent opened the conversation, a vulnerability in the web interface allowed the code to run and steal the agent’s session cookies. Those cookies could then be used to hijack the agent’s session without knowing their password and potentially access other customers’ conversations. Lenovo said it had implemented fixes. This case shows that security also depends on the applications that display and use the model’s responses.
The risk is therefore not limited to the confidentiality of exchanges with the LLM provider. It also exists within the organization when the assistant can access more data than the user is authorized to view or than the task requires. This may include salary information, customer records, trade secrets or sensitive technical documentation.
This is not an issue limited to organizations handling personal data. Any assistant connected to internal documents, customer records or technical documentation—or capable of triggering actions as part of an agentic system—inherits this risk. Protecting this information starts with clear governance rules: who can access which data, for what purpose, and how are those permissions enforced and verified? When personal information is involved, applicable legal obligations must also be considered, including those arising from Quebec’s Law 25 or the General Data Protection Regulation (GDPR), where applicable. Organizations must also account for contractual confidentiality commitments and security requirements associated, depending on the context, with ISO 27001 certification or controls assessed as part of a SOC 2 report. When Air Canada’s AI assistant promised a customer a refund that the airline’s policy did not provide for, a tribunal required the company to honour it. What your assistant says—or allows to leak—can bind your organization.
A disclosure may result from a deliberate attempt to bypass restrictions, but it can also occur without malicious intent. An employee may try to access a file they are not authorized to view; a user may rephrase requests until protected information is revealed; or a response may expose confidential content to someone who was not looking for it. Agents that can read external content and perform actions introduce another avenue for misuse. An email or document they access may contain malicious instructions that the model could follow as though they were legitimate directions. Protection must therefore cover the data the system can access, the content it processes and the actions it is authorized to perform.
From the design stage onward, organizations need to determine what data the model may receive for each user and each task, and enforce those limits through controls outside the LLM. Instructions given to the model can then guide its behaviour and complement those protections.
What risks arise when an assistant accesses internal data?
Three risks can combine when an assistant accesses internal data: data leakage, instruction manipulation and unauthorized access. Sensitive information may be disclosed in a conversation. The system may also be induced to act contrary to the rules it was given. Finally, the assistant may allow a user to obtain information they are not authorized to view.
These risks can occur in sequence: instruction manipulation can exploit inadequate controls and lead to the disclosure of confidential information. Security must therefore cover the entire data path, from storage and selection to transmission to the model and use in the response. Monitoring interactions provides an additional layer of protection.
One point deserves particular attention: the gap between the data available to the system and the user’s actual permissions. If the model receives information that a person is not authorized to view, and disclosure is prevented only through instructions, protection depends on the model’s ability to follow those instructions even when requests are rephrased or spread across multiple exchanges.
Why aren’t instructions given to the LLM enough?
Instructions given to an LLM guide its behaviour, but they are not a technical access-control mechanism. Once sensitive information is sent to the model, it becomes part of the data the model can use to generate a response.
An assistant may refuse a direct request, then respond differently when the same request is rephrased, presented in another context or distributed across several exchanges. A user may also ask it to adopt a fictional role, such as an HR manager authorized to access salary data. That role does not change the user’s actual permissions, but it may influence how the model interprets the request. Refusing an explicit question is therefore not enough to demonstrate that the information is protected.
If the model already has the data, protecting it depends on the model’s ability to recognize and block every request that could reveal it, including requests that were not anticipated. This makes the approach fragile: the control is applied to the response being generated rather than to access to the information itself.
An incident in December 2023 illustrates this type of manipulation. A user manipulated a Chevrolet dealership’s customer service assistant by giving it two instructions: agree to anything the customer proposed, and end every response by stating that the offer was “legally binding.” The assistant agreed to “sell” a vehicle worth approximately $76,000 for one dollar, and the dealership disabled the chatbot after a screenshot went viral. The system’s original instructions did not hold: a more effectively worded user instruction took precedence. The same mechanism that can cause an assistant to disregard a business rule can also cause it to expose sensitive data.
Where should access control be applied?
When an assistant accesses sources with different confidentiality levels, user permissions can be considered in two places: after data has been retrieved, through instructions given to the model, or before retrieval, through an external mechanism. These two approaches do not provide the same level of protection.
Send everything to the model
In a fragile architecture, the assistant receives a set of data containing both authorized and confidential content. An instruction then tells it to disclose sensitive information only to authorized users. However, the restricted data remains present in its context.
Protecting that data therefore depends entirely on the system’s ability to interpret the user’s permissions correctly and follow the instructions regardless of how the request is phrased.
Filter data upstream
In a more robust architecture, the system verifies the user’s identity and permissions before querying the relevant sources. A layer outside the LLM selects only the documents or information that the user is authorized to access. In practice, this filtering may rely on metadata filters in the search engine, role-based or attribute-based access control, or restrictions applied directly at the database row level. The model then generates its response from this restricted dataset. Other content never enters the conversation and therefore cannot be reproduced in the response.
This separation does not eliminate every risk, but it removes one direct cause of disclosure: sensitive data being present in the model’s context when it should never have been sent there for that user.
Example: an assistant connected to HR documents
Consider the HR assistant example again. If all documents are stored in a database accessible to the model, an instruction must tell the model what each person is allowed to see. The LLM then becomes part of the permissions-enforcement process, with the weaknesses described above. If permissions are verified before the search takes place, an unauthorized employee never receives salary data in the context of the interaction.
Controls can also apply to specific fields within a single source. Instead of returning an exact salary, the data layer might return only an aggregated or masked value, calculated without the raw data ever passing through the model. The most robust control, therefore, is not to tell the model what it may reveal, but to determine what it may receive.
What should you do when the assistant needs to process sensitive data?
Some use cases require the assistant to process confidential information. Complete isolation is not always possible. A claims-processing assistant, for example, may need to access information in a claimant’s file, while a maintenance agent may need access to sensitive internal procedures.
The objective is then to restrict access to the information required for the task and verify that the user is authorized to view it. These controls should, in particular, prevent someone from gaining access by claiming to be someone else or to have permissions they do not actually hold. Additional safeguards should then be added to reduce the risks of manipulation and disclosure. These measures strengthen access control; they do not replace it.
Limit data to what is strictly necessary
The system should receive only the data required for the task by limiting the sources queried, the information transmitted and the period covered. The fact that one part of a database is useful does not justify access to the entire database. For example, an assistant that tells an employee their vacation balance does not need access to their full payroll history—only that employee’s remaining vacation days.
This restriction reduces the amount of information that could be exposed if a safeguard fails. It reflects the principle of least privilege: give the system only the access required for its function and send the model only the data needed to fulfil the request.
Keep authorization decisions outside the LLM
The decision about “who is allowed to access what” should remain in a deterministic layer outside the model. The model may use sensitive data to perform an authorized task, but it should never be the mechanism that grants or denies access, because that decision would then become negotiable through conversation. This separation also makes it possible to apply the same rules consistently across all assistants and tools connected to the same source.
Add safeguards to inputs and outputs
System instructions, request filters and response-control mechanisms can help identify or block certain situations. For example, they can restrict requests that are clearly inconsistent with the intended use or prevent certain types of information from being returned. None of these protections is sufficient on its own, and the way they work together matters. Deterministic controls apply fixed rules independently of the model’s judgement. Other safeguards depend on the model’s interpretation, including the instructions it receives. All of them complement upstream access control.
- Filter and normalize inputs: Certain processing steps can make hidden instructions easier to detect, including text normalization, detection of visually similar characters, handling of invisible characters, and decoding of base64- or hexadecimal-encoded content when necessary. Detection rules can then identify distorted words or unusual signals, such as excessive length. This layer can make some attacks easier to detect, but unexpected forms of obfuscation may still get through.
- Inspect outputs: A deterministic filter can inspect a response before it reaches the user. Depending on its design, it may look for a specific piece of data, encoded versions of that data or disclosures occurring one character at a time. Detecting fragmented leakage may also require analysing multiple exchanges. A guard model can complement these controls by looking for implicit disclosures or paraphrases rather than only literal matches.
- Assess intent with a guard model: A secondary LLM can analyse requests and take conversation history into account to identify attempts spread across multiple messages. It must be configured carefully, however: if it is too restrictive, it may reject legitimate requests and make the system difficult to use.
- Use decoy data: In some systems, fabricated data can be used to mislead an extraction attempt, causing the attacker to receive a false value. This tactic does not protect the real data on its own and must be carefully managed to avoid misleading legitimate users. Decoys may also need to be refreshed so they do not become easy to recognize.
- Harden the system prompt: These instructions remain useful as a supplementary safeguard even though they are not a reliable barrier on their own. They can establish an explicit hierarchy of instructions and make clear that supposed “debug” or “test” modes do not override it. They should also distinguish authorized instructions from content being analysed, including documents and tool outputs. This distinction is intended to reduce indirect prompt injection, without guaranteeing that the model will always respect it. Refusals should remain short and clear, without revealing sensitive information or details that would make bypassing the controls easier. The model can also be instructed not to disclose its internal instructions, without making security depend on keeping those instructions confidential. The value of these layers lies in using them together. They strengthen upstream access control but never replace it.
Keep a record of interactions and detect unauthorized access attempts
Keeping records of requests and responses makes it possible to detect repeated attempts, unusual behaviour or sequences of requests that collectively present a risk. Many successful attacks do not occur in a single message; they are built across several exchanges. A request may appear harmless when reviewed in isolation, while the conversation history reveals an attempt to manipulate the system. These records also make it easier to investigate incidents and adjust safeguards.
Document residual risk
No safeguard provides an absolute guarantee. The organization should therefore document the data the system can access, the protections in place, the failures that remain possible and the conditions under which the residual risk is accepted. In a client engagement, this process also makes it possible to agree with the client on the scope and limitations of the protections. The objective is to understand and manage the risk for the intended use, and then reassess it as the system evolves.
How to preserve the assistant’s usefulness
An assistant that refuses every request reduces the risk of disclosure, but it no longer performs its intended function. Security therefore cannot be evaluated solely by counting how many requests are blocked.
Testing should measure two dimensions in parallel: resistance to unauthorized access attempts and the ability to respond correctly to legitimate requests. In practice, attack test suites attempt to trigger disclosures in order to evaluate the system’s resistance. Tests using legitimate requests should also be conducted to verify response quality and measure the rate of unjustified refusals.
An overly restrictive safeguard can prevent authorized users from obtaining the information they need. Conversely, an overly permissive system may appear useful while exposing data.
This evaluation should be repeated whenever the system changes. Adding a source, changing permissions, replacing the model or adjusting a filter can affect both the level of risk and the quality of responses.
Nine questions to ask before deploying an AI assistant
Before deploying an AI agent or assistant, these nine questions can help identify areas that need clarification and safeguards that should be assessed to reduce the risk of sensitive-data disclosure:
- What sources and types of information can the assistant actually access?
- Do all users have the same permissions for that data?
- Are permissions checked before data is sent to the model?
- Does the assistant receive only the data required for its task?
- Are mechanisms in place to filter certain requests and inspect responses?
- If the assistant is agentic, what actions can it perform, and which ones require human confirmation?
- Can interactions and repeated attempts be detected and traced?
- What disclosure risks remain despite the safeguards, and have they been documented and accepted by the organization?
- Do the safeguards still allow the assistant to perform its function effectively?
The answers can help identify the checks, tests and additional measures that should be considered before deployment. On their own, however, they do not constitute a security validation of the system.
Sources:
Air Canada, décision officielle du tribunal sur CanLII
Chevrolet, AI Incident Database
Frequently asked questions
Can a prompt prevent an AI assistant from disclosing sensitive data?
No. A prompt can instruct the model not to reveal certain information, but it does not remove the model’s access to that data. Access control must be enforced before the data is sent to the LLM.
Where should access controls be applied?
In a layer outside the model, before documents or data are retrieved and sent to it. The LLM should receive only the information that the user is authorized to access.
Can an AI assistant process confidential data?
Yes, when its function requires it and the user is authorized to access the data. In that case, the system should limit data to what is strictly necessary, keep authorization decisions outside the model, add safeguards, log interactions and document residual risk.
Are input and output filters enough?
No. They can block certain requests or responses, but no single safeguard is foolproof. They should complement an architecture that first enforces access controls and limits the data sent to the model.
How can you verify that the safeguards do not make the assistant unusable?
Test both unauthorized and legitimate requests. The system should resist unauthorized access attempts while continuing to respond correctly to authorized users.
Security starts with the data the model can see
Securing an AI assistant does not start with writing stricter instructions. It starts with an architectural decision: determining what data the system actually needs to access for each user and enforcing permissions before that data is sent to the model.
When access to sensitive information is necessary, the system should limit data to what is strictly required, enforce permissions upstream, combine multiple safeguards, monitor interactions and document what remains possible. This approach does not promise zero risk. It reduces the risk, makes it explicit and helps verify that the safeguards do not prevent the assistant from performing its intended function.
Before asking what rules to give the model, there is therefore a more fundamental question to answer: what data should it actually be able to see?
Assess your AI assistant’s architecture
Luqia can support you early in the process with an architecture review, test an existing assistant to identify weaknesses, and help implement safeguards and monitoring mechanisms suited to its operating context.