How a SaaS platform passed its first AI-in-scope audit without a scramble
The AI features were finally in scope for SOC 2 and the enterprise procurement questionnaire kept getting longer. Here is what changed when the evidence became a byproduct of the systems running, not a document assembled the week before.
5 min readStallwart
Where the work was breaking
The security team had passed SOC 2 twice on the same playbook: a spreadsheet inventory, a folder of policies, and a two-week scramble with screenshots before each audit window. The renewal year was different. Two AI features had shipped, the auditor had signalled they were in scope, an enterprise customer had sent a forty-question AI-specific questionnaire, and the board had asked whether the company was ready for ISO/IEC 42001 next year. The playbook that had worked twice was about to stop working.
The specific gap was evidence. The team could describe what the AI features did, in general, in a policy document. They could not, on demand, tell an auditor what a specific model version had been running on a specific date, what data it had seen, what decisions it had influenced, who had approved the change, or what stopped it from doing something it should not. The absence of that evidence, not the presence of AI, was the finding waiting to happen.
That is the specific bind teams end up in. Governance done at the policy layer alone reads as governance to an auditor who has not seen much AI, and reads as theatre to an auditor who has. Once AI is in scope, the auditor is going to ask for the evidence trail, and the only defensible answer is one that already exists.
What the system does instead
Sillage was pointed at the two AI features and stood up the governance layer as a byproduct of running them, not a project alongside them. A live inventory of every model in use updates as systems ship, so there is no gap between what the team believes is running and what is actually running. Each system carries a plain-language written basis for how it decides and what it is not permitted to decide, kept current in the same repo as the code.
Every high-stakes decision routes to a human by design rather than by luck. Runtime controls enforce policy at the moment of the decision, so a violation is prevented rather than caught after. Inputs, outputs, approvals, and overrides are logged continuously, retained, and queryable. Any automated action is reversible, and every system and every control has a named owner.
The auditor's questions become queries against evidence that already exists. What version was running on this date. What was the accuracy on the evaluation set that quarter. Which decisions were human-reviewed and which were fully automated. How was override used and by whom. The answers are produced in minutes because the record is a byproduct of the system, not a document reconstructed after the request.
Why this survives the questionnaire too
The enterprise procurement questionnaire and the auditor are asking for the same underlying thing in different vocabularies: an evidence trail that already exists, per system, in a form that can be produced on demand. When the governance layer is real, the same evidence answers both audiences, and the same answers hold up when the next auditor arrives with a slightly different vocabulary (ISO/IEC 42001 today, the EU AI Act's higher-risk obligations tomorrow).
That is the whole return on investment on getting governance into the system layer rather than the policy layer. A control that produces evidence as a byproduct is answered once and holds for years. A control that lives in a policy document has to be re-evidenced every audit cycle, and the effort scales linearly with the number of AI features shipped.
What changed for the team
The audit was answered from the log, not from a screenshot inbox. The enterprise questionnaire that used to consume a security engineer for a week was answered in hours because most of the questions were already covered by evidence the system was generating anyway. Legal stopped drafting bespoke language per customer because the same governance narrative now covered the same questions across customers.
The board question about ISO/IEC 42001 stopped being a project to start and became a scope conversation about what to certify against. Nothing new had to be built; the underlying evidence was already the shape that certification asks for. That is the outcome of putting governance in the system layer rather than the policy layer, and it is the specific reason Sillage exists.
What actually changes
- The audit answered itself. Model inventory, decision basis, approvals, and overrides were produced from the log in minutes rather than reconstructed from screenshots and memory across two weeks.
- The enterprise AI questionnaire stopped consuming a week per customer. The same evidence trail answered SOC 2, procurement, and forward-looking ISO/IEC 42001 and EU AI Act questions from one source.
- Governance stopped being a policy layer and became a system layer. Runtime controls prevent violations at the moment of the decision, not after, and every automated action is reversible.
We publish numbers once a customer has verified them. Nothing here yet, which is the honest answer.
Questions this raises
- How do you prepare for an AI-in-scope SOC 2 audit?
- By making the evidence trail a byproduct of the AI systems running, not a document assembled the week of the audit. A live model inventory, a written basis for how each system decides, human oversight for high-stakes decisions, runtime policy enforcement, continuous logging of inputs, outputs, approvals, and overrides, and rollback for every automated action are what an experienced auditor asks for once AI is in scope.
- What does ISO/IEC 42001 require for an AI management system?
- A governed, repeatable way of deciding what AI you deploy, how you assess its risks, who is accountable, and how you review it over time. It rewards evidence that already exists rather than a memo written before the review, which is the same underlying requirement as the AI-in-scope portions of SOC 2 and the higher-risk provisions of the EU AI Act.
- How should we answer enterprise AI vendor questionnaires?
- From the same evidence trail the auditor asks for, not from bespoke narrative language written per customer. Once the governance layer generates evidence as a byproduct of the AI systems running, the answers to the procurement questionnaire come out of the same source in a fraction of the time, and stay consistent across customers.
- Does AI governance need to be a project or can it be a byproduct?
- A byproduct is the only kind that survives an audit. Governance assembled the week of a review is a snapshot, not a control, and an experienced auditor can tell the difference. The evidence trail has to be generated continuously by the systems in production so it already exists when a regulator, customer, or board asks.
- How much AI compliance work is needed before the EU AI Act applies?
- That depends on the use case, because the EU AI Act is risk-tiered. Limited and minimal-risk uses carry light transparency duties; higher-risk uses require documentation, risk management, human oversight, logging, and traceability. The practical move is to classify each AI use early and map it to the obligations that tier actually triggers, so the compliance work is scoped to what applies.
Recognize this in your own operation?
Bring us the version of it happening in your business and we will tell you which part a system can take over.
Book a call