← Projects
Agent Autopsy — Forensic Debugger for Failed AI Agent Runs
Product demo
Problem
Most AI agent failures are silent — the run completes, the status code is green, and the result is still wrong. Observability platforms show you the trace, but never the cause of death.
Business Impact
Teams ship agents that fail silently and only hear about it from angry users. Agent Autopsy turns an opaque transcript into a one-line cause of death, so engineers stop guessing and start fixing — without shipping private traces to a third-party SaaS.
System Approach
- Next.js 16 App Router app with the forensic engine in lib/autopsy.ts
- scan + diagnose API routes: heuristic classifier first, optional LLM pass second
- Five explicit failure signatures detected from transcript structure, not vibes
- Optional local Ollama (qwen3:14b) pathologist's note for root cause
- 6 Vitest tests covering the signature classifier
Key Decisions & Trade-offs
- Heuristics-first for an instant verdict — the LLM is an optional deepening, not the critical path
- Local LLM over cloud — agent traces are sensitive and stay on the machine
- Five named signatures over a black-box score — engineers want a cause they can act on
- Zero runtime dependencies in the engine — rebuilt from a v1 single-file prototype with no shortcuts
Current Status
Public and open source (MIT) at github.com/rohitguta2432/agent-autopsy. Heuristic verdicts plus optional local qwen3:14b diagnosis working; 6 Vitest tests passing. The 15-second demo diagnoses the app's own sample corpse on camera.
Roadmap
- Add more failure signatures (silent truncation, tool-arg drift)
- Framework adapters to import LangChain / CrewAI transcripts directly
- Shareable report links for team triage
What I'd Improve Next
- A hosted demo with a bundled small model so no local setup is needed
- Batch autopsy mode to run inside CI on every failed agent run
Explore More
- MyFinancial — Personal Financial Advisor — Financial planning in India is fragmented across banks, insurance, and tax documents.
- PropCheck — AI Property Trust Score for India — Indian property buyers lose lakhs to fraudulent listings on Magicbricks, 99acres, Housing.
- StellarMIND — Chat-to-SQL with RAG — Business users need to query databases without knowing SQL.
- ClinicAI — WhatsApp AI Clinic Assistant — India has 12 lakh+ small clinics running on phone calls and paper diaries.
- MicroItinerary — AI Travel Planner — Travel apps optimize for proximity and ratings.
- SanatanApp — Hindu Devotional App — Devotional users in India juggle 5+ separate apps for Chalisa, Gita, Aarti, Ramayan, and Mahabharat.
- SynFlow — Enterprise Intelligence Platform — Private deal networks rely on manual introductions and spreadsheets.
- FinBaby (Jama) — Personal Finance Tracker — Indian middle-class families track expenses across UPI apps, bank statements, and paper notebooks.
- RetailOS — Multi-Tenant Retail SaaS — Indian kirana stores and small retailers use paper registers or basic billing software with no inventory tracking, no GST compliance, and no offline support.
- TripHive — Offline-First Collaborative Trip Planner — Group trip planning is fragmented across WhatsApp, Google Docs, Maps, Splitwise, and email.
- ScamRakshak — On-Device AI Scam Detector — Indians lose thousands of crores annually to digital scams via WhatsApp, SMS, and social media.
- PaisaGuard — Family Budget Survival App — Middle-class families worldwide track expenses inconsistently — UPI apps show transactions but don't enforce budgets.
- rohitraj.tech — Engineering work is often invisible.
- tinyvoice — Fine-Tune a Model in Your Own Voice in an Afternoon — Training your own language model sounds like a PhD job that needs a GPU cluster, so most developers never try.
- snap3d — One Photo In, an Editable 3D Model Out — A photo shows you one side of an object; the other five sides are a guess.
- Agentic OS — Force-Directed Map of a 387-Skill Claude Code Setup — An agent setup grows one plugin at a time until it has hundreds of skills, and nobody can see what it actually contains.
- claude-autodev — Autonomous 8-Stage Dev Pipeline for Claude Code — Agentic coding tools stop at a diff.
- KisanSathi — Six AI Farm Experts, Keyless and Local — Farmer-facing AI tools die at the API-key step, answer in English, and hand back generic advice with no live numbers behind it.
- MarginChef — AI Agent That Finds a Restaurant's Margin Leaks — Restaurants run on 3-5% net margins while food costs are up roughly 35% since 2019, and most owners never see per-dish economics.
- Quorum — Deep-Research Agent Swarm with a Shared GraphRAG Brain — Most AI research tools are one model in a loop.
- RegexForge — Plain English to a Regex You Can Trust — Regex is not hard because the syntax is exotic.
- Skillet — Turn Any Docs Page Into an Installable Claude Code Skill — Agent skills are the fastest way to teach a coding agent a new tool, but writing a good SKILL.
- Ladle — Open-Source Prep Forecasting and Food-Cost Leak Detection — Restaurants throw away 4-10% of the food they buy before it reaches a plate, and most owners find out at month end from a food-cost percentage that moved the wrong way, with no idea which ingredient did it.
- Casita — Design Your Home in 3D in the Browser — Home design tools want an install, a login, or a CAD background.
- VoxelForge — A Voxel Sandbox Engine Built Properly — Voxel sandboxes are the demo everyone builds with a coding agent right now, and almost all of them are a throwaway single HTML file that hitches on chunk generation, ships a texture pack, and cannot be read or extended.
- prompt-ocean — Type a Sea, Watch It Exist — Generative interfaces usually hand a model the wheel and hope.
- HEXAPOD — Inverse Kinematics and Gait Simulator — Hexapod gait and leg IK are usually explained with equations and a video, or hidden inside a robotics library.
- avatar-sync — Real-Time Face and Hand Tracking in One HTML File — Face and hand tracking demos come wrapped in a build step, a server, and usually an API key — which puts a wall in front of anyone who just wants to see what the models actually output before building on them.
- Reliability & Production Readiness — Load testing, observability, and API contracts.
- Open Source Repos — Browse the source code behind these projects.