SOC 2 usually arrives as a sales problem wearing an engineering costume. A deal is contingent on it, the date is already agreed, and someone forwards a spreadsheet of a few hundred questions to the person who runs the servers.
Then the panic-buying starts — a compliance platform, a scanning tool, a vendor promising a badge in six weeks. Some of that helps. But the reason SOC 2 feels impossible to scope is simpler than it looks: SOC 2 doesn't tell you what to build. There is no list of required services, no approved architecture, no minimum instance type. It's not that kind of standard.
That's genuinely confusing, and it's also the key to doing it without gold-plating.
The thing to understand first
SOC 2 doesn't audit your infrastructure. It audits whether you do what you say you do.
You describe your controls. The auditor checks whether the evidence supports that description, consistently, over a period. That's it. Which means an unremarkable, well-run, boring environment passes comfortably, and a sophisticated environment where the practice doesn't match the document does not.
This is why "which tools are SOC 2 compliant?" is the wrong question and gets sold expensive answers. No tool is SOC 2 compliant. Your process is, or isn't. A tool can produce evidence for a control you already run. It cannot produce the control.
The practical consequence: the most common failure isn't a missing control. It's a control you genuinely perform, that you can't prove you performed on the third Tuesday of March.
What auditors actually ask your infrastructure to demonstrate
Underneath the questionnaire, the infrastructure questions collapse into a handful of things:
Who can reach production, and how do you know?
Access has to be granted deliberately, reviewed periodically, and — the part that catches people — removed when someone leaves. Offboarding evidence is where a lot of findings land.
How does a change get to production?
Reviewed, approved, traceable. If deploys happen from a laptop with credentials nobody can enumerate, there's no story to tell here.
Would you notice something going wrong?
Monitoring, alerting, and log retention — plus evidence that alerts reach a human who acts. Silent monitoring is a finding, not a control.
Can you recover, and have you checked?
Backups are the easy half. Restore testing is the half that's usually theoretical, and it's the half that gets asked about.
Are your systems consistently configured?
Not perfectly hardened — consistently. An auditor wants to know that what's true of one server is true of the fleet, and that you can show it.
Do you know what you have?
An asset inventory that's current. Hard to fake, hard to maintain by hand, and quietly the thing that makes every other control provable.
Notice what's absent: nothing about your language, your cloud provider, your architecture, or whether you use containers. These are questions about discipline, not about technology.
Why it's hard anyway
If the requirements are that reasonable, why does SOC 2 hurt?
Because "consistently, across the fleet, with evidence, over a period" is a much higher bar than "we do that." Most teams do all six things above — on the servers they remember, in the way each engineer prefers, without a record. Every one of those is a genuine control and an audit finding at the same time.
The date is the trap. A Type II report observes your controls over a window — commonly 3–12 months. You cannot retroactively have been doing something. Teams that start when the deal lands discover that the earliest possible report date is months after the work is finished, not after it's started. If a deal depends on SOC 2, the audit window is the deadline, not the audit.
The order that works
Scope it down, aggressively
SOC 2 covers the systems in scope. A smaller, clearly bounded production environment is dramatically cheaper to certify than a sprawling one. Reducing scope is the highest-leverage hour you'll spend on this.
Make the fleet consistent before you make it compliant
Configuration automation first. If every server is built the same way from the same definition, a control demonstrated on one is demonstrated on all of them. Without this, every control is an audit of individual servers, forever.
Fix access and offboarding
Cheap, unglamorous, and the source of a disproportionate share of findings. Who has production access, why, and what happened when the last person left?
Put change control in the pipeline, not in a document
If the pipeline enforces review and records what shipped, your change-management evidence generates itself. If it's a policy people follow by hand, you'll be assembling evidence manually forever.
Test a restore, and keep the record
Not the backup — the restore. Do it once properly and you've answered a question that makes most teams uncomfortable.
Write down what you actually do
Last, not first. Policies written before the practice exists describe a fiction, and the gap between them is exactly what the auditor is employed to find.
The sequence matters. Steps 2 and 4 mean the evidence is a by-product of running the system properly. Skip them and every subsequent audit period is a manual evidence-gathering project.
What this looked like on a 70-server fleet
We implemented the infrastructure supporting SOC 2 for a global cybersecurity enterprise — a fleet of more than 70 servers across the UK, US and Australia. Worth noting who the client was: a cybersecurity company, being audited on security. There was no room for a compliance story that didn't hold up.
The core of the work was baseline automation with Chef and Ubuntu Pro, producing CIS-compliant environments from a definition rather than from a runbook, plus a CI pipeline that streamlined releases for the engineering team.
Two things worth taking from it:
The compliance work and the good-engineering work were the same work. Nothing was built for the auditor. Consistent, reproducible environments and a pipeline that records what shipped are what you'd want anyway; SOC 2 was the reason the business finally funded them.
It made the environment cheaper, not more expensive. The same audit found 20% of resource requirements weren't needed — removed with no impact to uptime or SLAs. That's not a coincidence. Once every server is built from a definition, the ones that shouldn't exist have nowhere to hide. Knowing what you have is a compliance control and a cost lever at the same time.
What's real and what's panic
Real: configuration automation, access review and offboarding, change control in the pipeline, restore testing, an accurate inventory, alerting that reaches a human.
Panic: buying a compliance platform before you have controls for it to evidence; hardening beyond CIS baseline because it feels safer; re-architecting for the audit; a policy library that describes a company you aren't yet.
The panic list isn't useless — it's just premature. A compliance platform is genuinely good at collecting evidence for controls that exist, and genuinely useless at inventing them.
Frequently asked questions
None specifically. SOC 2 doesn't mandate a stack, a cloud provider, or an architecture — it audits whether you do what you say you do, with evidence, over a period. In practice the infrastructure questions come down to access control, change management, monitoring, backup and restore, consistent configuration, and an accurate inventory.
Longer than the remediation work, because a Type II report observes your controls over a window of typically 3–12 months. You can't retroactively have been doing something, so the audit window — not the audit — is the real deadline. If a deal depends on it, start before the deal does.
Not a missing control — an unprovable one. Teams genuinely perform access reviews, restores and change approvals, but can't demonstrate they did it on a specific date. Evidence that's a by-product of the system beats evidence that someone has to assemble.
Eventually, and not first. These tools collect evidence for controls you already run; they can't create the controls. Buying one before the controls exist gives you a dashboard of things you aren't doing.
Not necessarily. The work it forces — consistent, reproducible, documented environments — is what lets you see what you're running. On a 70+ server fleet we certified, the same review cut resource requirements 20% with no impact to uptime or SLAs.
Related but different: ISO 27001 certifies a management system against a standard, SOC 2 is an auditor's report on your controls. The underlying infrastructure work overlaps heavily — do it once, and most of it serves both.
Get SOC 2-ready without gold-plating
Infrastructure that passes the audit because it's well run, not because it was built for one.
Related reading
- SOC 2 & ISO 27001 Compliance Audit — the engagement version
- DevOps Consulting — the infrastructure underneath
- How to cut AWS costs without cutting reliability — the same audit, the other outcome
- What a release should actually do — change control that evidences itself
- DevOps consulting for cybersecurity — what this looks like in a regulated sector