How Karya trains autonomous FDEs on top of an opinionated stack, and what it takes to make them safe enough to ship.
In draft
From Group-Drive to Unit-Drive
The productivity paradox of agentic software. A century-old lesson from American manufacturing, and the architecture decision every engineering organization now faces.
May 18, 2026
I.Electrification
In the 19th century, American factories relied on group-drive systems for power transmission. A single steam engine transmitted power via an intricate system of rotating shafts, belts, and pulleys. Factories were architected, manned, and capitalized around the needs of this engine.
In the 1890s, as factory owners began to adopt the electric motor, they retrofitted existing systems for simplicity. The improvements proved numerous: continuous power, quieter machines, reduced energy needs, improved speed control, and cooler working environments. On paper, they had achieved all the wins that electric motors had to offer.
In reality, the Total Factor Productivity (TFP) stagnated. The output of the factory was only marginally improved, while the capital input increased significantly (a new engine). It took 2 decades to realize electric motors facilitated the unit-drive which enabled independent machine placement. The result: modular layouts that adapted to rapidly changing manufacturing processes which beget exponentially compounding output.
This sort of overlaying of one technical system upon a preexisting stratum is not unusual during historical transitions from one technological paradigm to the next.
Paul A. David, 1990
II.Market
Organizations using coding agents to build software are experiencing this same productivity paradox. Most organizations today are replacing hand-written code with agent-generated code without modifying the factory: release processes, documentation, observability, local development, developer onboarding, etc.
They are bolting agents onto an architecture optimized for legacy technology.
Better coding agents / harnesses
Better memory recall
More context optimization
Cheaper VMs running existing Containers
Create more abstraction (hosted vibe coding)
We believe this is the wrong approach.
The illusion of immediate productivity makes these businesses attractive when in reality customers need to be retooled.
III. Agent-Centric Development
Some top-tier engineering organizations have figured this out, and have published their thoughts:
Agents must be infused with the business’ technical “taste”.
Organizational context / business processes must become visible to the agent.
Agent’s work must be verifiable in hermetic environments that operate concurrently.
New context must remain visible to both humans and agents.
Cultural shifts were required to begin adopting a workflow that’s truly AI native.
The result is that TFP skyrockets. Output compounds exponentially, while the cost is fixed: setting up a background agent.
This is summarized nicely by Stripe’s Alistair Gray:
As it turns out, parallelism, predictability, and isolation were also very desirable properties as well for Stripe engineers to be able to work most effectively. What’s good for humans is good for agents, and building on this infrastructural primitive paid dividends as a natural home for LLM agents.
Alistair Gray, Stripe
What’s implicitly buried in this thought leadership is the human intelligence behind it. This style of work requires what we canonically call a 10× engineer. A whole team of them.
Attracting this kind of talent is inherently impossible for the majority of firms. We believe we can address that gap.
IV. Where We Fit
Karya’s mission is to enable companies to build high-quality, vertically-integrated software autonomously.
It’s our belief that businesses should have access to truly autonomous background agents to build software in a safe, reliable, and maintainable way.
We do this by providing the opinionated tools, infrastructure, and processes (collectively, the stack) that enable any organization to achieve an end-to-end autonomous development workflow.
The benefits to our customers are numerous:
Network effects of shared infrastructure driving lower unit costs.
Optimizing output per-token value through improving code quality and imbuing taste.
Transparent workflows that benefit existing human engineers. (“what’s good for agents is good for humans”)
Accessibility for non-technical team members to contribute idiomatic code.
Product knowledge captured systemically for training.
Out of the box compliance readiness and heightened security posture.
Our version of meeting our customer where they are is through Forward Autonomously Deployed Engineers. An opinionated stack allows us to train and maintain highly proficient agents acting as FDEs to truly augment and upskill our customer’s workforce while helping continuously maintain their software.
V. The Future
Our vision is a world where the Ivory Tower is achievable and software is vertically integrated.
A business that adopts Karya’s technology can reasonably build any software that it needs bespoke. The decision to buy vs. build truly turns into a business decision surrounding alpha. We believe that our agents should be able to deeply understand the context of a business and offer maintainable software solutions to their problems without requiring third-party vendors.
Standardizing the stack also allows for a true network externality effect as we can observe and heal the tooling and infrastructure based on the collective issues faced by a larger cohort of businesses. Paul David referred to this as compatibility standardization.
VI. Our Positioning
Regarding the evolution of SoTA LLMs, Karya’s model positions us to benefit from advancements in frontier intelligence. A better underlying LLM means better service from our agents.
On adoption, there is a valid criticism regarding applicability to existing code bases. Those are a “known quantity” to an agent. Through standardizing what high quality output looks like however, we believe our agents can assist customers in migrating brownfield projects fully autonomously. The input is already constrained, we constrain the output.
The productivity ceiling, the four levers that raise it, and why the forward deployed engineers of the future won’t be human.
May 20, 2026
I. The Productivity Ceiling
Every new product brings the same inevitable question: how long until it’s ready to release? It’s a simple question on the surface, but at its heart sits a fundamental tension: speed of development vs. quality of development.
Think of the decision like a chart of productivity. Speed is on the y-axis and quality on the x-axis. The curve bows outward. You can have moderate amounts of both, but at the extremes the trade-off becomes punishing. Inherently, there is a maximum productivity that can be achieved without changing underlying variables: talent, tooling, process, or environment.
The productivity ceiling. Every organization picks a point along the curve: pure speed, pure quality, or somewhere in between. Agentic development pushes the ceiling outward.
II. Four Levers: A Golf Analogy
I find golf to be an easy analogy here. Think of your organization as a professional golfer and your productivity as golf score.
Talent: Better players have better scores. No sugarcoating this one: an organization is the sum of its parts, and talent is objective.
Tooling: Modern equipment makes the ball go farther and straighter. Frontier AI models and agentic development tooling have resulted in higher quality code faster — a true game changer.
Process: Professional golfers use a range finder to find distance, consult their caddy, and take a practice swing all before actually taking a shot. Software development requires careful thought, collaboration, and planning before any code is written.
Environment: Strategy and scores differ depending on the difficulty of the course, weather, etc. The composition of a codebase and the quality of a dev environment impact the quality of code written.
Historically, raising the ceiling of productivity meant pulling these levers one at a time. You hired a better engineer, adopted a new tool, tightened your process, or cleaned up your codebase. Each change was only marginally additive.
III. The Compounding Effect
When deployed effectively, agentic development changes that math. It collapses the tooling stack into something faster and more capable, absorbing process grunt work so that high-judgment tasks receive more attention. This in turn raises the effective skill of every engineer on the team and allows a small team to navigate, refactor, and maintain a codebase that previously would have demanded twice the headcount. The four levers can now be moved together, resulting in a compounding, rather than merely additive, effect.
IV. The Rise of the Forward Deployed Engineer
This then begs the question: how does an organization unlock agentic development? In response, the market for FDEs (Forward Deployed Engineers) has exploded. Organizations that lack the expertise are turning to industry experts who advise them on this new path.
FDEs install tooling and processes that optimize both speed and quality by leveraging the latest frontier AI models.
The FDE model works. Just look at the success of Palantir alongside the new enterprise AI services ventures from Anthropic and OpenAI. But is it optimized? It requires massive capital expenditure, and when the engagement ends, the expertise walks out the door. Still, the agentic development space continues to evolve rapidly. FDEs install tooling and processes optimized for today, but environments slowly degrade without the expertise to maintain them as best practices change and new code is written. The compounding effect is lost.
V. Autonomous Forward Deployed Engineers
So the next step is clear: FDEs must remain in place to avoid losing expertise or degrading the environment. The economics of keeping a human FDE on retainer indefinitely are out of reach for most organizations. This is where Autonomous Forward Deployed Engineers come in. An AFDE is an AI agent that lives inside your codebase and tooling, continuously maintaining and evolving your environment as the frontier shifts and your software grows.
Unlike their human counterparts, AFDEs don’t churn. They stay current with best practices automatically, integrate directly into your tooling and repo, and operate continuously, which means the compounding effect doesn’t decay the moment an engagement ends. And they do it at a price point that finally makes the math work.
Karya is building autonomous forward deployed engineers, making the compounding effects of agentic development accessible to every organization, not just the few that can afford to rent the expertise. The ceiling on productivity is still there. We just made it a lot cheaper to push through it.
How the modern software workflow strands context across six tools, and why agents inherit the wreckage.
August 18, 2026
I. The era of AI driven efficiency.(?)
Building software in the era of AI is an incredibly fun experience. What used to take weeks to months in agonizing planning sessions can now be condensed into a matter of days using a variety of shiny new tools. However, are we so sure that these tools are actually driving efficiency?
Take development of the most unglamorous feature in enterprise software: a data table. Filter the results, click a row to open the entity page, review the record, navigate back.
Here is how that screen actually got built in my past life. I suspect it will sound familiar.
I prototyped the interaction in Vercel's v0 until the flow felt right. I took my prompt to ChatGPT and decomposed it into requirements for a PRD. Designers replicated everything in Figma. Engineers reviewed the PRD and drew up engineering design documents based on their knowledge of our systems. Those became a set of Jira issues. At the end of the line, an agent does all of the implementation for you. It’s all done in a manner of days.
Sounds efficient, right? Let’s dig a little deeper…
II. Context Evaporated
Every handoff in that process dropped context the agent could use:
The prototype knew the intent. The v0 build encoded exactly how the interaction should feel: instant return, no refetch, state intact. The agent never sees it.
The PRD knew the why. The business rules and edge cases lived in a ChatGPT session and a doc three tools upstream. The agent gets two sentences.
The design doc knew the how. An engineer already reasoned about caching strategy and invalidation. The ticket carries the conclusion, but sees none of the reasoning.
Notice that no one in this story did anything wrong. Every step was diligent. The PRD was thorough, the design doc was thoughtful, the ticket was well written by the standards of tickets. The loss is structural. The workflow was designed to move work forward, not to move context forward, and agents can only act on the context that arrives.
III. The Smartest Teammate Outside the Room
Here is the strange part. The agent at the end of that line is the most knowledgeable engineer your organization has ever had access to. It has absorbed decades of data-fetching patterns, every flavor of state management and client-side routing, and every failure mode these systems have ever exhibited in the wild. It types faster than your whole team combined and never gets tired.
And yet think about how we treat it compared to any human engineer. An engineer sits in the planning meetings. They see the prototype, read the PRD, absorb the design doc, and soak up context through standups, code review, and a hundred Slack threads. By the time they touch the feature, they know why it exists and what "done" actually means.
The agent gets none of that. It joins the project at the very last step, receives a two-sentence ticket, and is expected to infer everything the room already knows. So not only is it lacking context, but it never had the opportunity to challenge any of the assumptions in the decision making process.
We hired the smartest teammate in the building, locked them out of every meeting, slid a sticky note under the door, and graded them on the result.
No organization would do that with their engineers, why are agents any different?
IV. A Seat in the Room
The market's instinct is to wait for a smarter teammate. Better agents, better models, better prompts. But no amount of talent compensates for being locked out of the room. So the fix is not a smarter agent. It is a seat in the room: a structure that delivers to the agent everything a trusted teammate would already know. The organizations getting real leverage from agents, the Ramps and Stripes of the world, converged on exactly this insight: agentic development is an infrastructure problem, not a tooling problem. That structure must do four things:
Carry intent to the agent. The prototype, the requirements, and the architectural reasoning must arrive with the task, not be re-derived from a two-sentence ticket. When product and engineering collaborate on one surface, the specification is the context, and there is nothing to drift.
Direct the work. Agents are versatile. For this reason, they are more difficult to use at enterprise scale. Complex tasks require complex solutions, and in a sea of possibility direction must be systematic. Golden paths, guardrails, and standing advisors on the hundreds of decisions that are made far before a ticket is ever created.
Verify the outcome, not the code. The review must shift from "does this diff look right" to "does this system behave correctly." Was any other dependency impacted? Will this handle the load of our existing users simultaneously navigating the page? Can this handle filtering across records if our database grows by 10x? Etc.
Capture the value. Every decision, constraint, and production signal must remain visible to both humans and agents, so the tenth table starts smarter than the first. The room's knowledge should accumulate, not evaporate.
Bolt-on tools cannot deliver this, because each one is another handoff. These properties only emerge when the environment is built as a whole.
V. Where We Fit
Karya ensures agents are always in the room. Product and engineering collaborate in one place, so the prototype, the requirements, and the reasoning live where the work happens, and intent flows to the agent instead of dying in a ticket. Every change is verified, approved by your engineers, and documented.
The tools of the AI era made us faster at every individual step and left the space between the steps untouched. That space is where the efficiency leaks out. Close it, and the smartest teammate you have finally gets to work the way the rest of the team always has: with full context, clear direction, and a verified definition of done. We believe every organization deserves that environment.
Uber, Ramp, and Stripe each spent tens of millions of dollars building an AI-SDLC in private. The common thread underneath all three is code standardization — and nobody has productized it.
August 30, 2026
I.Background
On Aug 27th 2026, Uber published an article detailing how they had begun to optimize the AI-SDLC (i.e. Software Factory) across roughly 5,000 engineers. Reading it and watching the accompanying presentation at AI Engineer reminded me eerily of a similar post just 2 months prior — Inspect at Scale from Ramp. Upon re-reading that, it just again felt similar to yet another article from Stripe about their Minions and what powers them.
We are now approximately 9 months into the new era of software engineering post-Opus 4.5. Coding agents are more than capable of literally every aspect of the SDLC, from triage, ideation, validation, architecting, to self-healing. Numerous new frontier models have emerged (e.g. GPT 5.6 Sol, Fable 5, Kimi K3, Grok 4.6), and somehow the learnings at these companies are still the world’s best kept secret. Despite these few examples, practically every non-SWE still asks: “Why can’t Claude Code just do it?”
These organizations have, despite being customers of Cognition, Factory, OpenAI, and Anthropic, still decided to invest tens of millions of dollars and thousands of man-hours to architect these solutions. The ROI makes sense, nearly 70% of code authored at these companies are now driven autonomously. Surely other companies can benefit from this; the question is how to productize the autonomous software factory.
II.Research
I wanted to approach the question with a research first approach. I leveraged Claude to index my own thoughts alongside a few hundred collated articles, op-eds, video transcripts, and anything else Autoresearch was capable of pulling up via Web Search relevant to the topic of AI-SDLC / Software Factory.
Fig 1. Tag-induced cluster structure of the Karya Brain vault. Nodes are wiki notes (n = 84), sized by degree; edges are resolved wikilinks (m = 387 undirected, 523 directed).
This data was output into a graph database, and the output clusters and their edges are visualized above. There were several themes that seemed to recur:
Agent Execution — Providing agents with all the necessary infrastructure to execute code reliably to validate their logic. (i.e. Uber DevBoxes, Inspect’s 30min refresh interval)
Excellence looks like MicroVM’s provisioned with all the dependencies necessary to execute logic with filesystem snapshots for quicker boots (< 5 seconds). These VMs are isolated away from production dependencies with strong network level isolation.
Workflow Orchestration — An opinionated set of workflows meant to guide agents through a “standard” workflow. These were custom built for each “shape” of software (web app, mobile app). (i.e. Stripe Blueprints)
Excellence looks like an opinionated set of curated skills provided to agents along with deterministic executable scripts (lints, formatters, static code analyzers, unit tests) that offered rapid validation of output logic.
These teams also spent considerable time shaping channels for non-technical team members to contribute (e.g. UIs like Inspect, Slack integrations) to facilitate more usage.
Context Management — Like any agentic system, context is king. For software systems, this explicitly translates to the code. This is the heart of the system, and the “substrate” itself.
Excellence looks like a well architected code base leveraging more static analysis and deterministic guardrails than agent skills which are inherently stochastic. Preference for strongly typed languages and convention-as-code.
Dynamic Verification — This layer is entirely about running agent generated code in a “production-like” environment and validating its performance dynamically. This involves a deeper understanding of all the external dependencies of your software and accurately faking them (e.g. Digital Twin) or leveraging them (i.e. FLOCI / LocalStack / pgmem).
Excellence looks like leveraging in-memory variants of external dependencies like Redis, Postgres, ElasticSearch while leveraging comprehensive Digital Twins of stateful third party dependencies like Okta or PostHog.
Considerable investment should be made in technologies like database branching (i.e. Copy on Write) to regression test logic against real production data safely.
Agent Observability — This has dual meaning. We need to reason about tool calls, MCP usage, and KV cache-optimization for repeated workflows just as much as we need the agent to have telemetry about running services so it can account for all constraints.
Excellence looks like standardized OTEL as an invariant across all deployed code with (S)PII redaction to ensure agents have total visibility into CPU & Mem utilization, happy path metrics, and traces across multiple dependencies to isolate bottlenecks.
Platform & IDP — Every piece of mission critical software has real SLAs with financial repercussions for failure. This must obviously be codified so that the right humans in the loop act as approval gates for mission critical software.
My first reaction to this is the same as Alistair Gray’s at Stripe: what’s good for humans is good for agents too. None of this is novel, these are all the same best practices prior to AI. If I were to imagine myself as an engineer back at Google in 2017, the engineering organization had:
Clearly defined style guides enforced through static analyzers (i.e. TAP Presubmit) that ran in less than 1-2 minutes that provided guidance on code structure and architecture.
Systems Integration Tests (i.e. Faker & Scuba) that leveraged inbuilt inversion of control in our application framework, Boq, to allow us to hermetically test logic prior to wasting compute on the cloud on our local machines. These systems sent back screenshots of the running application for golden diff tests.
Automated CI/CD that created preview environments once the first two checks passed to setup a staging environment for our change leveraging real production settings. If our code executed in amd64 on our local Intel macs, they were testing the workload on arm64 on Borg in a Google Data Center.
Eye of Sauron was auto-integrated into every piece of software automatically for dynamic analysis. Engineers could reason about CPU flame graphs, mem consumption, thread locking, etc. for all software without writing any custom logic.
Warm caches allow us to iterate on code when a bug was detected without waiting another 45 minutes for a fresh build and re-deploy. In the open source world, technologies like Hot-Module-Reloading also solve the same problem.
Ariane launch review combined with CODEOWNERS ensured that all changes had human approval as a gating mechanism. High visibility changes needed to clear eng leads, legal, PM, and even product design.
This real workflow back from 2017 was enabled by the standardization and infrastructure built by over 2,000 engineers at Google. Now, it seems, that similar investment at the major firms discussed above is paying dividends for a true AI-SDLC.
Analogically, the US Interstate system had to exist with well maintained asphalt, policing, gas pumps, and standardized signs so that drivers could go from point A to B quickly. The advent of self driving cars only leveraged the same infrastructure to drive faster while reducing human input.
III. Code Standardization
The common thread in the breakdown above is code. It’s the atomic unit of work in the AI-SDLC, and every aspect of the remaining infrastructure relies around it. At Google, Boq, the internal app development framework, afforded a standardization where it was easy to reason statically about:
External Dependencies (network boundaries for a program)
Compute Requirements (the code hinted at the cardinality of processed data)
Deployment Regime (all Boq programs compiled the same way for the ARM Axion processors they were deployed on)
Backwards Compatibility (Google3’s perforce monorepo along with Blaze’s build graph allowed CI/CD runners to evaluate affected systems and halt changes that could break production due to schema drifts).
I would wager that every single one of the companies above have invested heavily in code standardization over the last decade to ensure that their AI-SDLC can operate efficiently. Why?
Standardization drives a network effect within the organization.
The more programs that are similarly shaped, the more collective operational experience the organization can draw on. That is why Boq is a requirement at Google. It means that any engineer (now, any agent) can assess the impact of a change, reason about business logic, understand compute requirements, and even deploy or revert code in production through the same mechanisms as 1,000 different applications. It also means that as novel systems/techniques are discovered, they can be applied to the benefit of every team in unison.
Consider Stripe. They invested heavily in a type-checker for Ruby called Sorbet long before their AI-SDLC, Minions. That’s because they were running into loss of context due to Ruby’s qualities as a dynamically typed language across their 15 million line code base. This was fundamentally a human problem. 4 years later, they blogged about how they build Minions:
LLM agents are incredibly good at building software from scratch when there are relatively few constraints on a system. […] Humans must build sophisticated mental models to make effective changes in our repos, and enabling agents to develop the correct intuitions and use the correct tools within the confines of their context windows is challenging. […] Stripe has invested in developer productivity foundations that support our unique constraints at all stages in the development lifecycle—source control, environments, code generation, CI, and much more—and so our custom minion harness tightly integrates with that tooling.
Stripe
IV.Anti-Thesis
There are several companies now beginning to sell a means to build a software factory. Factory’s Software Factory, Overcut, Ona (now OpenAI), Harness, Cortex, and several more provide the tools to build-your-own workflow orchestrator backed by agent execution primitives (sandboxes) and varying degrees of context management tooling (e.g. source graphs, autowiki).
Ironically, the same innovations on frontier intelligence unlocked something far more valuable. What used to take 2,000 engineers to manage and maintain at Google (Boq) can now be reasoned about by just a handful armed with coding agents.
I believe this is where the real opportunity is. It’s less sexy, much deeper, but the benefits at scale have already been proven. And yet no one has productized this.
V.Opportunity
What if we focused on providing a developer productivity platform, recursively self-improved through observation at a scale larger than Google’s, Stripe’s, Ramp’s or Uber’s with industrial code standardization for the larger market?
Investing in a minimal set of abstractions (framework + compute) that offers maximum versatility in the shape of the software being deployed (i.e. web apps, mobile apps, APIs, data pipelines) and building the infrastructure around it to automate agent execution, context management, workflow orchestration, dynamic validation, observability, and governance would create a true Software Factory.
Leasing the factory to customers offers a true win-win. They can now begin to leverage developer productivity tools normally reserved for the largest Fortune 100 software companies and continuous maintenance and improvement of the infrastructure while leveraging AI to build software that is perpetually maintained autonomously. In parallel, by offering the same standardization to the broader market, the system can self-improve at a scale the world has not yet seen.
The key lies in making the right abstraction decisions. Historically, each decision would have to be evaluated manually and through experiments at scale which were expensive and slow. This is what coding agents now allow us to do but without the manual workforce. We can rapidly experiment numerous approaches focusing on the invariants we already know to work such as Kubernetes.
The applied version of this standardization would be the product to sell, which would enable customers to upskill their workforce 100x without the complications of configuring their own workflows, triaging why agents are stuck, or why their token costs are ballooning.
I think the future will be dominated by a few approaches shaped like the above versus the myriad point solutions that currently exist a-la-carte. If the world is moving towards focusing on outcomes not solutions, software engineering already has dozens of successful companies to look at to see what enables outcomes at scale.