Sandeep Panda
Blog

The software factory you can buy: FactoryKit, now Buildful

Ramp and Uber built software factories for their own engineers. We built one too, then turned it into something anyone can run for $10 a month.

By Sandeep Panda
10 min

tldr: A software factory is a pipeline where AI coding agents write and test the code and people review the pull requests. Ramp and Uber built theirs for their own engineers. We built one too, called it FactoryKit, and then turned it into Buildful, which anyone can use: describe a change, get a tested pull request on your GitHub repo. 1,000 tasks for $10 a month.

This post first described FactoryKit. FactoryKit is now Buildful, and this version is updated for what it does today (October 2026).

In July I wrote here about FactoryKit, the software factory we built so that Hashnode and Bug0 could keep shipping with a very small team. I'd write a task, an AI coding agent would pick it up in a cloud sandbox, and some minutes later a pull request came back with a recording of the change working in a real browser. I'd watch the recording, read the diff, and merge. By the end of July I'd mostly stopped opening my local development environment.

The Buildful run view for a small task, adding a contact email link to the footer: the prompt, the agent's narration and checks on the left; on the right, the opened pull request and the QA recording of the change in a browser
The Buildful run view for a small task, adding a contact email link to the footer: the prompt, the agent's narration and checks on the left; on the right, the opened pull request and the QA recording of the change in a browser

Back then I thought the people who needed this were engineering teams, and that most of them would buy one rather than build one. I was right that most people can't build one. I was wrong about who needed it most.

What is a software factory?

A software factory is a development pipeline where AI coding agents do the production work, implementing changes, running checks and testing the result in isolated environments, while people direct the work and review it through pull requests. You hand it a task and get back reviewed code, with evidence that it works.

The term is older than AI agents. Microsoft used "software factories" in the 2000s for process frameworks and code generation, and the US Department of Defense runs whole programs under the name. When people say software factory in 2026, they usually mean what Ramp and Uber built: agents produce the code and humans review it before it ships. Some say "AI software factory" to make clear which meaning they intend. That's the sense I use here.

How Ramp and Uber run their software factories

Ramp built a background coding agent called Inspect. Each session runs in a sandboxed VM on Modal with what an engineer would have locally, including Vite, Postgres and Temporal, so the agent can check its own work the way an engineer would. In January 2026, Ramp wrote that about 30% of pull requests merged to its frontend and backend repos were written by Inspect. A month later, Modal's case study put it at roughly half.

Uber built Minion, which Gergely Orosz describes as an internal background agent platform with monorepo access. An engineer gives it a prompt, and a few minutes later Slack says a pull request is ready to review. In March 2026, Uber's CTO, Praveen Neppalli Naga, said about 1,800 code changes a week were being written entirely by that agent, and that 95% of Uber's engineers use AI tools every month.

The two setups share a shape. The agent works away from the engineer's laptop, in an environment that looks like the real one. It runs the same checks a person would. A human reviews the result as a pull request. Both companies built all of this themselves, with their own engineers, for their own engineers.

That's the catch for everyone else. A ten-person startup can't spare engineers for a project like that, and a student or a freelancer can't do it at all. I wrote up what it takes to build one: seven parts that each look easy in a demo.

From coding assistants to background coding agents

The last few years of AI coding tools look like a ladder. Autocomplete suggested lines. Chat wrote functions you pasted in. Agents inside the editor, like Cursor and GitHub Copilot, edit files while you watch. Each step moved more work to the machine, and each one still needed you sitting there.

A background coding agent is the step where you can leave. It takes a task, starts an isolated environment, writes the code, runs your checks and comes back with a pull request. You review its output the way you'd review a teammate's.

This is also where vibe coding and agentic coding split. Vibe coding is prompting your way to software you never read. Agentic coding hands a defined task to an agent and puts a human review at the end. I wrote more about the difference in vibe coding vs agentic coding.

How a task runs in Buildful today

Here's what happens when you give Buildful a task. There's nothing to install. It all runs in the browser and in the cloud.

  1. Connect GitHub. Install the Buildful GitHub App and pick the repos it may work on.
  2. Describe the change. Plain words are fine. Pick one or more repos, and attach screenshots if they help. You can also start a task from Slack with /buildful or an @Buildful mention, or by adding a label you chose to a Linear issue.
  3. Buildful picks the model. It reads the task and picks the right open-weight model for it. The work runs in a private cloud sandbox with a copy of your code.
  4. Your checks run. Lint, types, tests: whatever your repo uses. If a check fails, the agent fixes it and runs it again, up to 3 attempts, then reviews its own diff.
  5. UI changes get tested in a real browser. The agent starts your app, uses the changed feature the way a person would, and records the session.
  6. You get a pull request per repo. Each one has a summary, the agent's self-review, what it verified, its caveats and the recording. Nothing reaches your main branch until you merge.
  7. Follow up if you need to. Reply on the task to ask a question or ask for a change. New commits land on the same pull request.

A few rules from the FactoryKit days stayed. The agent never runs git. It edits files, and Buildful commits, pushes and opens the pull requests. Each task gets its own sandbox, and the sandbox is deleted when the task ends. Our own API keys and GitHub tokens are added by the sandbox firewall as requests leave, so the agent only ever holds a placeholder. Your repo's environment variables do go into the sandbox, because your app needs them to run, so give Buildful staging or development credentials, never production ones.

A pull request opened by Buildful: what it verified, its caveats, and the QA recording in the pull request body, above the original task
A pull request opened by Buildful: what it verified, its caveats, and the QA recording in the pull request body, above the original task

The recording is the part I'd miss most. A green check tells you the code does what its tests expect. Watching the feature work in a browser, before you read the diff, tells you whether it does what you asked.

Why we stopped selling a factory

FactoryKit was going to be a software factory as a service for engineering teams. We did demos. We offered forward-deployed engineers who'd set it up inside a company's infrastructure, and self-hosting for teams that needed it. It looked like the obvious business.

Two things changed our minds. The first was the cost structure. Once the sandboxes, templates, browser testing and pull request plumbing exist, they cost about the same whether ten people use them or ten thousand. What grows with use is mostly model tokens. The second was who kept asking about it. Much of Hashnode's traffic comes from India and Africa, and the developers we hear from there are often students, freelancers and people building on their own. They want what Ramp's engineers have, and they can't pay \(100 or \)200 a month for an AI plan. Even Claude Pro at $20 comes close to ₹2,000 a month in India once forex and GST are added.

So we asked who else could use what we'd built, and rebuilt it around that answer. What Buildful sells now is tasks:

  • 1,000 tasks for $10 a month, which is about a cent a task.
  • Buildful picks the right open-weight model for each task: a fast one for a small fix, and the most capable one, with more time to think, for a hard task. Open-weight models cost far less to run than frontier ones, and that's what makes a cent a task possible.
  • Frontier models, whenever you choose. On the paid plans you can pick Claude or GPT yourself for a task and pay from prepaid credit, or run it on your own API key.
  • You can start a task from a Slack message or a labelled Linear issue, and the pull request goes back to the issue.
  • It learns from merged pull requests. After each task, Buildful proposes notes about your repo, and they become active when you merge. Up to 5 per repo go into future tasks.

There are no demos, no sales calls and no enterprise contracts. You sign in, connect GitHub and give it a task.

What it can't do yet

Buildful works with GitHub only. If your code lives on GitLab or Bitbucket, it can't open pull requests there yet.

Browser testing works best for web apps. Changes with no runnable UI, like a library or a config file, still run your checks, but there's nothing to record.

A vague task still gets you a vague pull request. The tasks that come back ready to merge read like a good ticket: what's wrong, where, and how you'll know it's fixed. That was true with FactoryKit on frontier models, and it's still true now.

Checkout is in US dollars only for now. Local prices are something we want, and they aren't built.

If those limits fit how you work, the plans are on the Buildful pricing page.

FAQs

What is a software factory?

A software factory is a development pipeline where AI coding agents implement, check and test changes in isolated environments while people direct the work and review it through pull requests. The term once meant process frameworks and code generation. Today it usually describes what Ramp and Uber built for their own engineers.

How is Buildful different from GitHub Copilot or Cursor?

Copilot and Cursor are best known for helping inside your editor while you work, and both now offer agents that work in the background too. Buildful only works in the background: you describe a change in the browser, Slack or Linear, and you get back a pull request. Every change runs your checks, every UI change is tested in a real browser and recorded, and you pay per task: 1,000 for $10 a month.

Is there an open source software factory?

You can build one from open-source parts. Coding agent CLIs like OpenCode and Codex are open source, and so is Playwright for driving a browser. The sandboxing, credential handling and the loop that ties them together are what you'd still have to build. Buildful itself isn't open source.

What happened to FactoryKit?

FactoryKit is now Buildful, at buildful.ai. Same people, the makers of Hashnode, and the same core loop: a task goes in, a tested pull request comes out. What changed is who it's for and how it's priced. The old /factorykit Slack command still works.

#ai#software-development#agentic-coding#developer-tools#testing

Discussion3

Add a comment

More writing

All writing