claude-autodev — Autonomous 8-Stage Dev Pipeline for Claude Code
Problem
Agentic coding tools stop at a diff. Nothing forces a spec to exist before implementation, nothing blocks a run that never wrote a test, and nothing catches a stage that quietly produced no artifact — so failures surface as a plausible-looking PR nobody can trust.
Business Impact
The gap between an AI-written diff and shippable software is spec, review, and test discipline — exactly the parts teams skip when the agent looks confident. claude-autodev makes those parts mandatory gates, so an unattended run either produces a reviewed, tested PR or stops with a diagnosis instead of a false green.
System Approach
- Eight stages, each a fresh headless claude -p session with zero shared context between them
- Artifact gates per stage — no artifact, no advance
- Isolated git worktree per run so concurrent runs never collide
- SQLite registry plus per-run events.jsonl as source of truth; HTTP+SSE server drives the dashboard read-only
- Bounded retries with self-fix, then park BLOCKED with a written diagnosis
- Daemon mode turns tracker issues into runs and issue labels into the queue
Key Decisions & Trade-offs
- No shared context between stages — an independent reviewer is worth more than a cheaper one
- Gate on artifacts, not on the model saying it finished
- Worktree isolation over branch switching so a stuck run never blocks the main checkout
- Server is a viewer, not a dependency — the pipeline completes with the dashboard closed
- Deploy stage is opt-in; merging and shipping should never be an accident
Current Status
Public and open source (MIT) at v0.2.0. Test suite runs on Linux, macOS, and Windows across Node 22 and 24 in CI on every push. Ships autodev selftest, which drives a fixture repo through all seven stages in about 30 seconds without spending model quota, and autodev doctor for preflight checks with a fix per failure.
Roadmap
- Multi-repo runs so one requirement can span services
- Pluggable review policies per repository
- Cost and duration budgets per run with hard stop
What I'd Improve Next
- Hosted dashboard mode for teams watching several runs at once
- Richer holdout scenario authoring so the test gate is harder to game
Explore More
- MyFinancial — Personal Financial Advisor — Financial planning in India is fragmented across banks, insurance, and tax documents.
- PropCheck — AI Property Trust Score for India — Indian property buyers lose lakhs to fraudulent listings on Magicbricks, 99acres, Housing.
- StellarMIND — Chat-to-SQL with RAG — Business users need to query databases without knowing SQL.
- ClinicAI — WhatsApp AI Clinic Assistant — India has 12 lakh+ small clinics running on phone calls and paper diaries.
- MicroItinerary — AI Travel Planner — Travel apps optimize for proximity and ratings.
- SanatanApp — Hindu Devotional App — Devotional users in India juggle 5+ separate apps for Chalisa, Gita, Aarti, Ramayan, and Mahabharat.
- SynFlow — Enterprise Intelligence Platform — Private deal networks rely on manual introductions and spreadsheets.
- FinBaby (Jama) — Personal Finance Tracker — Indian middle-class families track expenses across UPI apps, bank statements, and paper notebooks.
- RetailOS — Multi-Tenant Retail SaaS — Indian kirana stores and small retailers use paper registers or basic billing software with no inventory tracking, no GST compliance, and no offline support.
- TripHive — Offline-First Collaborative Trip Planner — Group trip planning is fragmented across WhatsApp, Google Docs, Maps, Splitwise, and email.
- ScamRakshak — On-Device AI Scam Detector — Indians lose thousands of crores annually to digital scams via WhatsApp, SMS, and social media.
- PaisaGuard — Family Budget Survival App — Middle-class families worldwide track expenses inconsistently — UPI apps show transactions but don't enforce budgets.
- rohitraj.tech — Engineering work is often invisible.
- Agent Autopsy — Forensic Debugger for Failed AI Agent Runs — Most AI agent failures are silent — the run completes, the status code is green, and the result is still wrong.
- tinyvoice — Fine-Tune a Model in Your Own Voice in an Afternoon — Training your own language model sounds like a PhD job that needs a GPU cluster, so most developers never try.
- snap3d — One Photo In, an Editable 3D Model Out — A photo shows you one side of an object; the other five sides are a guess.
- Agentic OS — Force-Directed Map of a 387-Skill Claude Code Setup — An agent setup grows one plugin at a time until it has hundreds of skills, and nobody can see what it actually contains.
- KisanSathi — Six AI Farm Experts, Keyless and Local — Farmer-facing AI tools die at the API-key step, answer in English, and hand back generic advice with no live numbers behind it.
- MarginChef — AI Agent That Finds a Restaurant's Margin Leaks — Restaurants run on 3-5% net margins while food costs are up roughly 35% since 2019, and most owners never see per-dish economics.
- Quorum — Deep-Research Agent Swarm with a Shared GraphRAG Brain — Most AI research tools are one model in a loop.
- RegexForge — Plain English to a Regex You Can Trust — Regex is not hard because the syntax is exotic.
- Skillet — Turn Any Docs Page Into an Installable Claude Code Skill — Agent skills are the fastest way to teach a coding agent a new tool, but writing a good SKILL.
- Ladle — Open-Source Prep Forecasting and Food-Cost Leak Detection — Restaurants throw away 4-10% of the food they buy before it reaches a plate, and most owners find out at month end from a food-cost percentage that moved the wrong way, with no idea which ingredient did it.
- Casita — Design Your Home in 3D in the Browser — Home design tools want an install, a login, or a CAD background.
- VoxelForge — A Voxel Sandbox Engine Built Properly — Voxel sandboxes are the demo everyone builds with a coding agent right now, and almost all of them are a throwaway single HTML file that hitches on chunk generation, ships a texture pack, and cannot be read or extended.
- prompt-ocean — Type a Sea, Watch It Exist — Generative interfaces usually hand a model the wheel and hope.
- HEXAPOD — Inverse Kinematics and Gait Simulator — Hexapod gait and leg IK are usually explained with equations and a video, or hidden inside a robotics library.
- avatar-sync — Real-Time Face and Hand Tracking in One HTML File — Face and hand tracking demos come wrapped in a build step, a server, and usually an API key — which puts a wall in front of anyone who just wants to see what the models actually output before building on them.
- Reliability & Production Readiness — Load testing, observability, and API contracts.
- Open Source Repos — Browse the source code behind these projects.