It's an exciting time to work in San Francisco. The city is alive with startups again, its freeway billboards advertise services that seem incomprehensible to most of the populace, and we are able to take a shocking number of our meetings with other AI companies in person. Several of our customers are near the Salesforce Tower, and I often ask them to meet me for a walk around the urban park on top of the trans-bay terminal. This is not an original idea; it is full of pairs doing their weekly 1:1's, VCs taking pitch meetings and now, software engineers one-shotting entire features of their product by sending a Slack message to Cursor while enjoying a $12 iced matcha.
Here in the 49 square miles of SF, or more specifically SoMa and a bit of the Financial District, companies like Anthropic, OpenAI, Cursor, Cognition, LangChain, Mercor, Vercel and dozens of smaller startups are living in the future of software development. Humans are writing little to no code and instead, software engineers are becoming something more akin to highly technical product managers, coaches, and architects as they describe what they want from their coding agents in more and more abstract ways.
This future isn't without challenges, especially for a CISO who is trying to live up to responsibilities defined decades ago while supporting such a wild shift in development. First, code volume is exploding. Companies embracing AI development can ship orders of magnitude more code than they did two years ago, as each IC engineer is now really a manager of a team of agents. Second, development is leaving the desktop, but not all at once. Engineers now kick off work directly from Slack or a Linear, and cloud agents spin up containers and write the code somewhere outside of the already porous security perimeter CorpSec teams have defined with EDR, Zero Trust networking, and MDM.
The third immediate challenge comes from the key conflict between this new software delivery model and the venerable assurance mechanism security teams and auditors have relied upon for years: code review. Human verification has become a serious bottleneck, and many companies have built mechanisms to automatically merge AI generated code without human review or approval before those mechanisms have become standardized. At Corridor, we're already starting to merge code without human review.
Finally, this is all happening at the same time that AI is getting very good at finding vulnerabilities in existing code. That is, the efficiencies gained from these new development models are an exciting solution to the immediate need to squash thousands of bugs, but it also means that autonomous development models need to apply to decades of hand-built, legacy code and can't only be used to blank-sheet projects.
In short: nearly every security control the industry built between 2002 and 2024 naturally assumed that a human wrote the code and a human could review it. Those assumptions are long past accurate, and the secure development practices we built over the last several decades will not survive the next two years, as the "coding from the rooftop" model spreads from SF across the country and the world (hopefully with cheaper coffee).
This post is about how we replace the traditional SDLC with something that gives us even stronger security assurances and that can operate at the 10-50x/developer scale of code generation needed as Enterprises adopt automated software development.
Six Levels of Software Development Automation
First, we need to define a standardized taxonomy of how software responsibilities are evolving from humans to AI. Dan Shapiro's blog post comparing AI development to the 6 levels of autonomy in self driving remains the best framework I've seen for easy categorization (although you can tell a human wrote much of his post because there is an off-by-one error).
- Level 0 — Autocomplete: AI acts as a reference tool or smarter search engine; the developer manually writes all the code, turning to AI only for specific snippets or answers.
- Level 1 — Coding Agent Assistance: AI writes boilerplate or unimportant code; the developer prompts for specific functions but immediately reviews and integrates everything it produces.
- Level 2 — Coding Agent Pair Programming: An interactive pair-programming partnership where the developer and AI trade off control, with the human reviewing code as it's generated in real time.
- Level 3 — Coding Agent Review: AI generates the majority of the codebase; the developer reviews everything it does, becoming the verification bottleneck before work can progress.
- Level 4 — Autonomous Coding Agents: AI runs unattended for long stretches on complex tasks; the human trusts the system's self-checks and only inspects the final feature much later.
- Level 5 — Software Factory: The engineer manages goals and the system rather than the code — providing plain-English descriptions while the AI defines the implementation, writes code and tests, fixes bugs, and ships.
As Shapiro points out, teams operating at L5 today are tiny (often less than a dozen people) and what they're producing is, in his words, nearly unbelievable. Small, agile teams are often where development practice shifts start before being adopted up the chain to the largest Enterprises. As organizations with deeper societal responsibilities climb this ladder, the software security field will need to keep up.
The Modern Era of Software Security (2002–2024)
If you dig into contemporary software security processes, you will find that they are built on years of failures, reactions and standardization, like the ancient ruins of Troy.
Modern application security started with the Trustworthy Computing memo Bill Gates sent to every Microsoft employee on January 15, 2002. Battered by Code Red, Nimda, and a never-ending parade of serious malware outbreaks, Gates paused all software development, declared security the company's highest priority and set the bar at computing "as available, reliable and secure as electricity, water services and telephony." Microsoft invested heavily in building their own team but also helped fund the burgeoning industry of consultants and tooling providers. This initial push by Microsoft gave us a new kind of formalized threat modeling, a Security Development Lifecycle tied to Microsoft's classic waterfall development model, developer security training at scale, compiler-enforced replacement of dangerous functions and, in 2003, the first Patch Tuesday.
The core components of the Microsoft SDL's, such as threat modeling at design time, secure coding standards during implementation, static analysis before ship, penetration testing after, became the skeleton of every "secure SDLC" for decades. OWASP gave the field a shared vocabulary for the emerging field of web security research with the first Top 10 in 2003. Adam Shostack and the STRIDE school made threat modeling a repeatable design-phase ritual rather than an art form. Gary McGraw's BSIMM work (circa 2008) let CISOs benchmark the maturity of their programs against peers.
The security industry rapidly shifted during this time, away from network protections and towards securing code. Static analysis vendors commercialized the "scan" step that Microsoft had to build internally. Dynamic testing tools sold the "verify" step, and the widespread adoption of bug bounties led to new economic models for both researchers and bounty platforms. After the massive Heartbleed coordinated-response effort demonstrated the risk of relying upon a small number of libraries for critical functions, software composition analysis became its own category. The move to the cloud and CI/CD gave us DevSecOps, and the scanners moved from quarterly engagements into the pull request.
At each step, compliance regimes have tried to capture the innately artisanal aspects of software security work. PCI DSS (2004) wrote secure development and code review requirements into a commercial mandate. SOC 2 and ISO 27001 require change management, usually enforced via human code review. After the SolarWinds incident, the U.S. government adopted new policies on software security and composition: Executive Order 14028 (2021), the NIST Secure Software Development Framework (SP 800-218), SBOM requirements, and attestation forms for anyone selling software to an agency. The EU's Cyber Resilience Act extended the same logic to any product with digital elements sold in Europe.
The Common Thread: Human Effort
If you look at the classic SDLC, you see a number of built-in assumptions around how software is created:
- Threat modeling assumes a design phase, where senior engineers or architects think through the design of a distributed system and document it before almost any software engineering is done.
- Static analysis assumes a scan-and-triage cadence measured in days or weeks, with a focus on fixing flaws after they are merged into main and perhaps even shipped.
- Code review assumes a qualified human reads every change and is willing to point out flaws. This is a critical assumption that has rarely held up except in the most expensive and sensitive software projects.
- Change management assumes approvals map to people or well-defined roles. Credentials, permissions, and audit trails all are tied to human identities and, ultimately, human accountability.
We can interrogate these assumptions against the Shapiro automation levels discussed earlier. As an organization reaches Level 2 they mostly hold: a human still applies every change, so the human gate still exists, even if the code's provenance is getting murky. At L3 the SDLC becomes heavily strained: the human is still in the loop, but the loop is spinning faster than any human can meaningfully review, and "approved" starts to mean "skimmed". At L4 the SDLC model is intentionally and selectively bypassed, as routine changes ship on the say-so of an automated reviewer, and humans see only what gets routed to them.
At Level 5 this model completely collapses. There is no human reviewer. There is no secure developer training. There is no triage meeting. An AI factory turns specs into software, and every control that assumes a human in the middle now has to be replaced with something that gives us equivalent levels of assurance with much lower latency and higher throughput, to deal with the massive increase in KLOC that comes from such levels of automation.
SADL: the Secure Agentic Development Lifecycle
While AI is better than humans at writing secure code in some circumstances, our experience tells us that coding agents' general lack of knowledge of business processes, architecture, and overall context create many more opportunities for complex flaws that emerge from AI-generated code. They also write so much more code. Clearly, we must figure out ways to continue the work that started in 2002 as we move towards complete automation of software development.
The team here at Corridor would like to propose a new set of security controls that match the moment. Not all of them are feasible to deploy at all organizations today, and few are commercially available, but in putting together this list it is clear to us that all of these controls are well within the technical capabilities of current LLMs and are necessary to establish trust in automated software development.
We believe a good name for the successor to the SDLC for the agentic age should be SADL, the Secure Agentic Development Lifecycle. As members of the community, we have no desire to trademark or otherwise own this acronym, and hope to encourage its open use by the community and our competitors.
Architecture and threat modeling are a separate, human-led activity, decoupled from coding, often by another team. The developer's intent, architecture, and security requirements live in tickets/design docs/wiki, not in the editing context.
Intent is expressed ad hoc in prompts; architecture and threat model are not systematically captured per change. Risk: the prompt conveys feature intent but omits architectural context, trust boundaries, and security requirements. Might use a SECURITY.md file.
Intent emerges conversationally as the pairing session unfolds; architecture and security context live in repo context files (e.g. CLAUDE.md / SECURITY.md) the tool reads. Threat modeling is still human-led, but the AI can flag risky design choices in-session.
The developer's intent must be captured explicitly as the task is set — goals carry architectural context and security acceptance criteria. A lightweight per-feature threat model is still human-authored; the agent can draft architecture/threat-model notes from the stated intent for human review.
Intent capture and threat modeling become harness steps: a planner/security sub-agent elicits and records the developer's intent, derives architecture and a threat model from the spec, and humans sign off on high-risk designs. Security/compliance requirements are encoded as machine-readable policy the harness enforces.
Architecture and threat model must be generated and enforced automatically from the captured intent/spec at trigger time. Deterministic risk classification (does this touch authz / PII / money / infra?) drives routing — eliciting and confirming developer intent is the main point where a human re-enters, via escalation.
To help organize our thoughts, we have simplified secure development controls and practices into ten categories. Seven of them map to the standard stages of development:
- Architecture and threat modeling - Capturing the developer's intent, the system's architecture, and its trust boundaries so that security requirements exist before code does, whether they're written by a human architect or derived from a spec by a planner agent.
- Secure-generation guardrails - In-the-loop controls that catch secrets, insecure patterns, and policy or compliance violations at the moment code is generated, before a change is ever accepted or committed.
- Security analysis - Discovering and remediating vulnerabilities in code, dependencies, and infrastructure configuration, with findings prioritized by reachability and delivered fast enough to act on within the development loop.
- Automated testing and verification - The test suites, evaluators, and quality gates that prove a change actually works and behaves as specified; since tests are how agents self-correct, suite quality directly bounds how much autonomy is safe.
- Code review and approval - Deciding whether a change is acceptable to merge, evolving from a human reading every diff to an automated reviewer with risk-based routing that reserves human attention for high-blast-radius changes.
- Change authorization and release governance - The policy layer that determines who (or what) may merge and deploy a change, and under what conditions (ex branch protection and human sign-off at one end, policy-as-code with progressive rollout at the other).
- Runtime protection and observability - Monitoring shipped software for anomalies and maintaining the ability to detect, roll back, and, at full autonomy, pause the factory when something agent-built misbehaves in production.
Three of the categories cut across these stages to provide additional visibility and control over AI agents and their output:
- Identity, access, and agent permissions - Giving agents their own scoped, least-privilege identities and sandboxed execution environments rather than letting them borrow their operator's credentials, because over-broad scope is how agents delete data or leak secrets.
- Provenance, audit, and traceability - Maintaining a tamper-evident record connecting every change back to the spec, prompt, model, tools, and reviewer that produced it, so decisions no human watched can be reconstructed after an incident.
- Feedback loops and autonomy governance - Continuously measuring agent output quality (escape rate, revert rate, security findings at review) and using quantitative gates to decide when a class of tasks earns more autonomy and when it gets demoted back to human review.
We then mapped these ten categories against the six levels of automation, from no automation at all to completely autonomous software factories.
Here are some themes that emerge as we climb the automation ladder while attempting to maintain the same levels of assurance:
Threat modeling stops being a meeting. At L0 it's a separate, human-led activity, decoupled from coding, often done by another team entirely. By L3, intent has to be captured explicitly as the task is set, because the agent will faithfully build whatever the prompt says, including the missing trust boundaries. At L4 it becomes a harness step: a planner sub-agent elicits intent, derives the architecture and threat model from the spec, and humans sign off only on high-risk designs. At L5 the threat model must be generated and enforced automatically at trigger time, with deterministic risk classification (i.e. does this change touch AAA, PII, money, or infrastructure?) driving to infrequent human escalation.
Guardrails move from advisory to blocking. At L1 guardrails are editor conveniences: linters and on-save secret scanners. At L3, where an agent can run bash and touch many files, secrets and unsafe data handling have to be intercepted in real time. At L5 guardrails are mandatory and blocking, with auto-remediation in-session, for the simple reason that code will be committed without additional human review. Compliance obligations have to be defined ahead of time and enforced deterministically as code, instead of non-deterministically via testing.
Reviews invert. The trajectory of code review is the clearest illustration of the whole shift. L1: a human reviews 100% of code. L3: a human still reviews everything, with AI-assisted summarization making the diff mountain survivable. L4: review becomes selective and risk-routed. An automated reviewer handles routine changes while criticality routing forces human eyes onto high-blast-radius ones (AAA, database, IaC, payments). L5: no human PR approval by default. An automated reviewer must match or exceed the capability of human review for in-scope changes, with a routing layer that tags each change and attaches deterministic consequences per tag.
We have included three categories of controls that are critical to maintain human control over this process. These replace processes that exist in companies between teams or individuals:
Agent identity. Agents can no longer borrow their operator's credentials. L4 and L5 demand per-agent (eventually per-run, ephemeral) identities, least-privilege tool and credential scopes, just-in-time secrets, egress controls, and hard sandboxing. As Agents become responsible for production deployments and troubleshooting, these identities become highly privileged and should be divided up carefully between agents of differing responsibilities. Just as with human employees.
Provenance. With humans out of the loop, it becomes critical to be able to "git blame" back to the spec, prompt, or guardrail that caused a failure. L5 requires tamper-evident, end-to-end provenance sufficient to reconstruct, after an incident, the decisions that no person observed. This will be critical in understanding which of the many pieces of context entering a factory as raw material lead to the faulty part on the backend.
Autonomy governance. A new application of the discipline of continuous quality evaluations. Quantitative gates (ex. escape rate, revert rate, follow-up-PR rate, security findings at review) govern how much autonomy each class of agent or task has earned. This is critical to being able to avoid feedback loops that eventually create unmaintainable, untrusted code that no human understands well enough to fix.
Conclusion
At Corridor, we already provide visibility, control and guardrail enforcement for desktop, long-running and cloud agents, and we are rapidly building new features and products to support our customers as they climb the ladder towards fully automated factories. We are also making this climb ourselves, and as we learn from our own challenges and mistakes we are rapidly integrating the lessons we learn into our product so that others can benefit from the solutions we have built to prevent the same missteps.
This is an incredibly exciting time to build, and we look forward to sharing our learnings and progress with you. Look for more blog posts, videos and updates from us and please feel free to sign up to receive updates and invites to our events so we can learn from each other.