All posts

Thoughts on Software Factories

Uber, Ramp, and Stripe each spent tens of millions of dollars building an AI-SDLC in private. The common thread underneath all three is code standardization — and nobody has productized it.

I. Background

On Aug 27th 2026, Uber published an article detailing how they had begun to optimize the AI-SDLC (i.e. Software Factory) across roughly 5,000 engineers. Reading it and watching the accompanying presentation at AI Engineer reminded me eerily of a similar post just 2 months prior — Inspect at Scale from Ramp. Upon re-reading that, it just again felt similar to yet another article from Stripe about their Minions and what powers them.

We are now approximately 9 months into the new era of software engineering post-Opus 4.5. Coding agents are more than capable of literally every aspect of the SDLC, from triage, ideation, validation, architecting, to self-healing. Numerous new frontier models have emerged (e.g. GPT 5.6 Sol, Fable 5, Kimi K3, Grok 4.6), and somehow the learnings at these companies are still the world’s best kept secret. Despite these few examples, practically every non-SWE still asks: “Why can’t Claude Code just do it?”

These organizations have, despite being customers of Cognition, Factory, OpenAI, and Anthropic, still decided to invest tens of millions of dollars and thousands of man-hours to architect these solutions. The ROI makes sense, nearly 70% of code authored at these companies are now driven autonomously. Surely other companies can benefit from this; the question is how to productize the autonomous software factory.

II. Research

I wanted to approach the question with a research first approach. I leveraged Claude to index my own thoughts alongside a few hundred collated articles, op-eds, video transcripts, and anything else Autoresearch was capable of pulling up via Web Search relevant to the topic of AI-SDLC / Software Factory.

Tag-induced cluster structure of the Karya Brain vault: six labelled clusters — Agent Execution, Context Management, Cost & Observability, Dynamic Verification, Platform & IDP, and Workflow Orchestration — drawn as a node-link diagram.

Fig 1. Tag-induced cluster structure of the Karya Brain vault. Nodes are wiki notes (n = 84), sized by degree; edges are resolved wikilinks (m = 387 undirected, 523 directed).

This data was output into a graph database, and the output clusters and their edges are visualized above. There were several themes that seemed to recur:

  1. Agent Execution — Providing agents with all the necessary infrastructure to execute code reliably to validate their logic. (i.e. Uber DevBoxes, Inspect’s 30min refresh interval)
    1. Excellence looks like MicroVM’s provisioned with all the dependencies necessary to execute logic with filesystem snapshots for quicker boots (< 5 seconds). These VMs are isolated away from production dependencies with strong network level isolation.
  2. Workflow Orchestration — An opinionated set of workflows meant to guide agents through a “standard” workflow. These were custom built for each “shape” of software (web app, mobile app). (i.e. Stripe Blueprints)
    1. Excellence looks like an opinionated set of curated skills provided to agents along with deterministic executable scripts (lints, formatters, static code analyzers, unit tests) that offered rapid validation of output logic.
    2. These teams also spent considerable time shaping channels for non-technical team members to contribute (e.g. UIs like Inspect, Slack integrations) to facilitate more usage.
  3. Context Management — Like any agentic system, context is king. For software systems, this explicitly translates to the code. This is the heart of the system, and the “substrate” itself.
    1. Excellence looks like a well architected code base leveraging more static analysis and deterministic guardrails than agent skills which are inherently stochastic. Preference for strongly typed languages and convention-as-code.
  4. Dynamic Verification — This layer is entirely about running agent generated code in a “production-like” environment and validating its performance dynamically. This involves a deeper understanding of all the external dependencies of your software and accurately faking them (e.g. Digital Twin) or leveraging them (i.e. FLOCI / LocalStack / pgmem).
    1. Excellence looks like leveraging in-memory variants of external dependencies like Redis, Postgres, ElasticSearch while leveraging comprehensive Digital Twins of stateful third party dependencies like Okta or PostHog.
    2. Considerable investment should be made in technologies like database branching (i.e. Copy on Write) to regression test logic against real production data safely.
  5. Agent Observability — This has dual meaning. We need to reason about tool calls, MCP usage, and KV cache-optimization for repeated workflows just as much as we need the agent to have telemetry about running services so it can account for all constraints.
    1. Excellence looks like standardized OTEL as an invariant across all deployed code with (S)PII redaction to ensure agents have total visibility into CPU & Mem utilization, happy path metrics, and traces across multiple dependencies to isolate bottlenecks.
  6. Platform & IDP — Every piece of mission critical software has real SLAs with financial repercussions for failure. This must obviously be codified so that the right humans in the loop act as approval gates for mission critical software.

My first reaction to this is the same as Alistair Gray’s at Stripe: what’s good for humans is good for agents too. None of this is novel, these are all the same best practices prior to AI. If I were to imagine myself as an engineer back at Google in 2017, the engineering organization had:

  1. Clearly defined style guides enforced through static analyzers (i.e. TAP Presubmit) that ran in less than 1-2 minutes that provided guidance on code structure and architecture.
  2. Systems Integration Tests (i.e. Faker & Scuba) that leveraged inbuilt inversion of control in our application framework, Boq, to allow us to hermetically test logic prior to wasting compute on the cloud on our local machines. These systems sent back screenshots of the running application for golden diff tests.
  3. Automated CI/CD that created preview environments once the first two checks passed to setup a staging environment for our change leveraging real production settings. If our code executed in amd64 on our local Intel macs, they were testing the workload on arm64 on Borg in a Google Data Center.
  4. Eye of Sauron was auto-integrated into every piece of software automatically for dynamic analysis. Engineers could reason about CPU flame graphs, mem consumption, thread locking, etc. for all software without writing any custom logic.
  5. Warm caches allow us to iterate on code when a bug was detected without waiting another 45 minutes for a fresh build and re-deploy. In the open source world, technologies like Hot-Module-Reloading also solve the same problem.
  6. Ariane launch review combined with CODEOWNERS ensured that all changes had human approval as a gating mechanism. High visibility changes needed to clear eng leads, legal, PM, and even product design.

This real workflow back from 2017 was enabled by the standardization and infrastructure built by over 2,000 engineers at Google. Now, it seems, that similar investment at the major firms discussed above is paying dividends for a true AI-SDLC.

Analogically, the US Interstate system had to exist with well maintained asphalt, policing, gas pumps, and standardized signs so that drivers could go from point A to B quickly. The advent of self driving cars only leveraged the same infrastructure to drive faster while reducing human input.

III. Code Standardization

The common thread in the breakdown above is code. It’s the atomic unit of work in the AI-SDLC, and every aspect of the remaining infrastructure relies around it. At Google, Boq, the internal app development framework, afforded a standardization where it was easy to reason statically about:

  1. External Dependencies (network boundaries for a program)
  2. Compute Requirements (the code hinted at the cardinality of processed data)
  3. Deployment Regime (all Boq programs compiled the same way for the ARM Axion processors they were deployed on)
  4. Backwards Compatibility (Google3’s perforce monorepo along with Blaze’s build graph allowed CI/CD runners to evaluate affected systems and halt changes that could break production due to schema drifts).

I would wager that every single one of the companies above have invested heavily in code standardization over the last decade to ensure that their AI-SDLC can operate efficiently. Why?

Standardization drives a network effect within the organization.

The more programs that are similarly shaped, the more collective operational experience the organization can draw on. That is why Boq is a requirement at Google. It means that any engineer (now, any agent) can assess the impact of a change, reason about business logic, understand compute requirements, and even deploy or revert code in production through the same mechanisms as 1,000 different applications. It also means that as novel systems/techniques are discovered, they can be applied to the benefit of every team in unison.

Consider Stripe. They invested heavily in a type-checker for Ruby called Sorbet long before their AI-SDLC, Minions. That’s because they were running into loss of context due to Ruby’s qualities as a dynamically typed language across their 15 million line code base. This was fundamentally a human problem. 4 years later, they blogged about how they build Minions:

LLM agents are incredibly good at building software from scratch when there are relatively few constraints on a system. […] Humans must build sophisticated mental models to make effective changes in our repos, and enabling agents to develop the correct intuitions and use the correct tools within the confines of their context windows is challenging. […] Stripe has invested in developer productivity foundations that support our unique constraints at all stages in the development lifecycle—source control, environments, code generation, CI, and much more—and so our custom minion harness tightly integrates with that tooling. Stripe

IV. Anti-Thesis

There are several companies now beginning to sell a means to build a software factory. Factory’s Software Factory, Overcut, Ona (now OpenAI), Harness, Cortex, and several more provide the tools to build-your-own workflow orchestrator backed by agent execution primitives (sandboxes) and varying degrees of context management tooling (e.g. source graphs, autowiki).

However, without investing in codebase standardization, teams will simply burn an ever increasing pile of compute and tokens solving and re-solving the same problems over and over with the varying flavors of implementation that a coding agent decides to take. What’s worse is then codifying all of these changes results in an over bloated AGENTS.md that must grow linearly with the surface area of the built software. Think building a skyscraper without an architect, you’re just reinforcing structure as you go.

Ironically, the same innovations on frontier intelligence unlocked something far more valuable. What used to take 2,000 engineers to manage and maintain at Google (Boq) can now be reasoned about by just a handful armed with coding agents.

I believe this is where the real opportunity is. It’s less sexy, much deeper, but the benefits at scale have already been proven. And yet no one has productized this.

V. Opportunity

What if we focused on providing a developer productivity platform, recursively self-improved through observation at a scale larger than Google’s, Stripe’s, Ramp’s or Uber’s with industrial code standardization for the larger market?

Investing in a minimal set of abstractions (framework + compute) that offers maximum versatility in the shape of the software being deployed (i.e. web apps, mobile apps, APIs, data pipelines) and building the infrastructure around it to automate agent execution, context management, workflow orchestration, dynamic validation, observability, and governance would create a true Software Factory.

Leasing the factory to customers offers a true win-win. They can now begin to leverage developer productivity tools normally reserved for the largest Fortune 100 software companies and continuous maintenance and improvement of the infrastructure while leveraging AI to build software that is perpetually maintained autonomously. In parallel, by offering the same standardization to the broader market, the system can self-improve at a scale the world has not yet seen.

The key lies in making the right abstraction decisions. Historically, each decision would have to be evaluated manually and through experiments at scale which were expensive and slow. This is what coding agents now allow us to do but without the manual workforce. We can rapidly experiment numerous approaches focusing on the invariants we already know to work such as Kubernetes.

The applied version of this standardization would be the product to sell, which would enable customers to upskill their workforce 100x without the complications of configuring their own workflows, triaging why agents are stuck, or why their token costs are ballooning.

I think the future will be dominated by a few approaches shaped like the above versus the myriad point solutions that currently exist a-la-carte. If the world is moving towards focusing on outcomes not solutions, software engineering already has dozens of successful companies to look at to see what enables outcomes at scale.

All postsBook a demo