Rohit delivered our MVP in 5 weeks — on budget and ahead of schedule. His architecture decisions saved us from rewriting everything when we scaled.
AI Engineering Foundations — Teach Your Team to Actually Use the AI Tools You Already Pay For
A two-week, hands-on programme on your own codebase: spec-driven development from zero, which model and effort level for which task in Claude Code and GitHub Copilot, the folder structure that gives every tool project context, and the token maths behind the bill.
The Problem
Your engineers have Copilot, Claude, Kiro, maybe OpenAI too. Leadership has seen a dozen proofs of concept. Yet nobody can say how much faster the team actually ships, because day-to-day usage is shallow: prompts typed straight into chat with no spec, the most expensive model on every task or the cheapest one on the hard ones, effort left on the default, no repository instructions, so every tool guesses at your conventions. Measuring productivity at this stage only measures the gap in basics. Close the fundamentals first — specs, model and effort choice, context files, review habits — and the gains become real enough to measure.
What You Get
- ✓Spec-driven development from zero: why specs beat ad-hoc prompting, and the full Spec Kit flow — constitution, specify, clarify, plan, checklist, tasks, analyze, implement, converge — run on a real feature from your backlog
- ✓Kiro specs side by side with Spec Kit: requirements in EARS notation, design, tasks, and steering files — so teams on either tool follow one process
- ✓Model and effort playbook for Claude Code and GitHub Copilot: which model and which effort level for spec writing, planning, implementation, debugging, review and support work
- ✓Repository context set up for your main repos: CLAUDE.md, AGENTS.md, .github/copilot-instructions.md, path-specific instructions and Kiro steering — the single biggest lever on output quality
- ✓Token and cost basics: context windows, input vs output vs thinking tokens, prompt caching, and a cost model per feature so spend is predictable
- ✓Separate tracks for Development (C++ and Node/React), QA (specs to test cases) and Support (docs and ticket retrieval)
- ✓Recorded demo, spec templates and a written playbook your team owns
- ✓Before-and-after skills check, plus named internal champions who carry the practice after the programme ends
Where This Fits
- 1FoundationsThis programme. The team learns specs, model and effort choice, context files and review habits.
- 2Team-wide tool setup, guardrails and a productivity baseline.
- 3AI systems shipped inside your product by an embedded engineer.
Demo: One Ticket, Spec to Pull Request
A real feature taken from a one-line ticket to a reviewed pull request with Spec Kit — and the model and effort level chosen at each step, with the reason.
Recording in progress — until it is up, the same walkthrough runs live on your own codebase in the free half-day demo.
- 0:00Install Spec Kit and run specify init --here in an existing repo
- 1:00/speckit-constitution — the project's non-negotiables
- 2:00/speckit-specify and /speckit-clarify on a real backlog ticket
- 3:30/speckit-plan with Opus at high effort, /speckit-tasks
- 5:00/speckit-implement on Sonnet at medium effort
- 6:30/speckit-converge, review the diff, open the PR
Spec-Driven Development, From Zero
Most teams use AI coding tools by typing a one-line request into chat. That works for a small edit and falls apart for a feature: the agent guesses the requirements, picks its own architecture, and nobody can review the result against anything. Spec-driven development fixes the order of work — what and why first, then how, then tasks, then code — and keeps each step in a file that humans review and the agent implements against.
The Spec Kit flow
Install once: uv tool install specify-cli, then, inside an existing repo, specify init --here --integration claude (or copilot). The steps run inside the agent's chat, not the terminal. Copilot and Claude Code both invoke them as /speckit-specify and so on; Spec Kit's docs write the same steps as /speckit.specify.
| Step | When | What it produces |
|---|---|---|
/speckit-constitution | Once per project | Principles every later step is checked against: testing rules, security rules, architecture constraints. |
/speckit-specify | Every feature | The what and the why: user stories and acceptance criteria. No tech stack here. |
/speckit-clarify | Quality gate | The agent asks up to five targeted questions and writes the answers back into spec.md. |
/speckit-plan | Every feature | The how: stack, architecture and constraints become plan.md and supporting design files. |
/speckit-checklist | Quality gate | “Unit tests for requirements” — checks the spec is complete and unambiguous before work is split. |
/speckit-tasks | Every feature | A dependency-ordered tasks.md: setup, foundations, one phase per user story, polish. |
/speckit-analyze | Quality gate | Read-only consistency check across spec, plan and tasks before any code is written. |
/speckit-implement | Every feature | Executes tasks.md phase by phase, respecting dependencies and parallel markers. |
/speckit-converge | Every feature | Compares the code with the spec; appends missing tasks. Repeat implement → converge until Converged. |
Small features take the short path: specify → plan → tasks → implement → converge. Production features add the three quality gates (clarify, checklist, analyze).
Kiro specs, for teams on Kiro
Kiro builds the same idea into the IDE. Each spec lives in .kiro/specs/<feature>/ as requirements.md (or bugfix.md), design.md and tasks.md. Acceptance criteria use EARS notation, which keeps them testable:
WHEN an auditor exports a report for a closed period THE SYSTEM SHALL produce a PDF signed with the firm's certificate AND record the export in the audit log with user, time and checksum.
What a good spec contains
| Weak spec | Strong spec |
|---|---|
| “Add export to reports.” | User story, who needs it and why, and 3–6 acceptance criteria a tester can check. |
| Mixes UI details, database choice and goals in one paragraph. | Spec says what and why; the stack and architecture go in the plan. |
| Silent on errors, permissions and limits. | Names edge cases: empty data, no permission, large files, time-outs. |
| No link to existing rules. | Inherits the constitution: test coverage, security and performance rules. |
The Folder Structure That Gives Every Tool Context
Output quality depends less on the model than on what the model knows about your repo. These files are how you tell it — once, in version control, for every engineer and every tool.
your-repo/
├── AGENTS.md ← shared instructions every agent reads
├── CLAUDE.md ← Claude Code project memory (Copilot reads it too)
├── .specify/ ← Spec Kit
│ ├── memory/constitution.md ← project principles, written once
│ ├── templates/ scripts/
│ └── feature.json ← which feature is active
├── specs/
│ └── 001-export-audit-report/ ← one folder per feature
│ ├── spec.md ← what & why
│ ├── plan.md ← how (+ research, data model, contracts)
│ ├── tasks.md ← ordered work items
│ └── checklists/requirements.md
├── .claude/ ← Claude Code
│ ├── skills/speckit-*/SKILL.md ← Spec Kit skills (specify init --integration claude)
│ ├── agents/ commands/
│ └── settings.json ← permissions, hooks, per-model effort
├── .github/ ← GitHub Copilot
│ ├── copilot-instructions.md ← repo-wide instructions
│ ├── instructions/cpp.instructions.md ← path-specific, applyTo: "src/**/*.cpp"
│ └── skills/speckit-*/SKILL.md ← Spec Kit skills (specify init --integration copilot)
└── .kiro/ ← Kiro
├── specs/export-audit-report/
│ ├── requirements.md ← user stories + EARS acceptance criteria
│ ├── design.md
│ └── tasks.md
└── steering/ ← Kiro's equivalent of instruction filesCopilot reads .github/copilot-instructions.md, path-specific .github/instructions/*.instructions.md files, the nearest AGENTS.md, and a root CLAUDE.md. Keep shared rules in AGENTS.md so a team using several tools maintains one source of truth.
Which Model and Which Effort, for Which Task
Two dials matter. The model sets the ceiling on capability and the price per token. The effort level sets how many tokens the model spends thinking, calling tools and writing — on Anthropic's current models, effort is the main control for depth and cost. The expensive mistake is running everything at one setting.
Claude Code
Switch model with /model opus, /model sonnet, /model haiku or /model fable; set effort with /effort high or claude --effort xhigh. The opusplan alias plans with Opus and implements with Sonnet — a good default for spec-driven work. Per-model effort can be pinned in .claude/settings.json under modelSettings, and a skill or subagent can set its own effort in frontmatter.
| Task | Model | Effort | Why |
|---|---|---|---|
| Constitution, spec, clarify | Opus 5.5 | high | Wrong requirements waste days; this is where depth pays. |
| Plan and task breakdown | Opus 5.5 (or opusplan) | high | Architecture choices and task ordering need reasoning. |
| Implement a well-specified task | Sonnet 5.5 | medium | The spec already carries the thinking; speed matters more. |
| Implement a hard or long task | Sonnet 5.5 or Opus 5.5 | high | Raise effort before switching model. |
| Hard C++ bug, legacy module, cross-cutting refactor | Fable 5.1 or Opus 5.5 | high → xhigh | Long reasoning chains across unfamiliar code. |
| Autonomous runs over ~30 minutes | Opus 5.5 or Fable 5.1 | xhigh | Anthropic's documented use case for xhigh. |
| Code review of AI-written changes | Opus 5.5 | high | A reviewer should be at least as capable as the author. |
| Small edits, summaries, ticket triage, subagents | Haiku 4.5 or Sonnet 5.5 | — / low | Fastest and cheapest; Haiku has no effort setting. |
GitHub Copilot
Choose the model in the chat input's model picker; for reasoning models, the arrow next to the model name opens a Thinking Effort menu (None, Low, Medium, High). Auto lets Copilot route by task complexity — fine for everyday chat, but pin the model for spec and review work so results are repeatable. The same Claude models are available in Copilot alongside OpenAI's GPT and Codex models, Gemini and others; which ones you see depends on your plan and your organisation's policy.
| Task | Model in Copilot | Thinking effort |
|---|---|---|
| Spec, plan, architecture questions | Claude Opus 5.5 or a top GPT model | High |
| Agent-mode implementation from tasks.md | Claude Sonnet 5.5 or a Codex model | Medium |
| Inline completions and quick chat | Auto or a fast model (Haiku 4.5, a mini model) | Low / None |
| Reviewing a pull request | Claude Opus 5.5 or equivalent | High |
Larger models use more of a Copilot plan's premium allowance, so the routing above is also a budget decision.
Effort levels explained
| Level | Use it for | Trade-off |
|---|---|---|
low | Simple, well-defined tasks; subagents; chat. | Fastest and cheapest; may under-think hard problems. |
medium | Well-specified agentic coding. Default on Opus 5.5 and in Claude Code on Sonnet 5.5. | The balance point for most implementation. |
high | Complex reasoning, difficult coding, specs and plans. API default on most models. | Spends what the task needs. |
xhigh | Long-running agentic work, roughly 30+ minutes. | Meaningfully more tokens than high. |
max | Frontier problems only. | Often large cost for small gains; can overthink. |
Rule of thumb: raise effort before switching to a bigger model, and lower it only where you have checked that quality holds.
Tokens and Cost, the Basics
- Token: the unit models read and bill in. On Claude's current tokenizer, 1M tokens is roughly 555,000 words.
- Context window: how much the model can hold at once — 1M tokens on Fable 5.1, Opus 5.5 and Sonnet 5.5; 200K on Haiku 4.5. Bigger is not free: every token in context is billed on every turn.
- Input vs output: output costs five times input on current Claude models, and thinking counts as output. That is why effort level moves the bill.
- Prompt caching: repeated context (instruction files, the spec, unchanged code) is read from cache at a fraction of the input price — 10% on most models, 5% on Opus 5.5, 2.5% on Fable 5.1.
- Batch: non-urgent jobs through the Batch API are 50% off.
Current Claude models
| Model | Best for | Input / output per 1M tokens | Context | Default effort |
|---|---|---|---|---|
| Claude Fable 5.1 | Hardest reasoning, long agentic runs | $10 / $50 | 1M | high |
| Claude Opus 5.5 | Specs, plans, agentic coding, review | $4 / $20 | 1M | medium |
| Claude Sonnet 5.5 | Everyday implementation | $2 / $10 | 1M | high (medium in Claude Code) |
| Claude Haiku 4.5 | Fast, cheap, high-volume | $1 / $5 | 200K | not supported |
API list prices, last verified 29 September 2026 against Anthropic's models overview. Seat-based plans (Claude Team and Enterprise, Copilot) bill differently; the relative costs still hold.
Worked example: one feature, two ways
A mid-sized feature: about 150K input and 40K output tokens for spec, plan and tasks, then about 2M input (80% cached) and 150K output for implementation.
| Approach | Spec + plan | Implementation | Total |
|---|---|---|---|
| Everything on Fable 5.1 | $3.50 | $11.90 | ≈ $15.40 |
| Opus 5.5 plans, Sonnet 5.5 implements | $1.40 | $2.62 | ≈ $4.02 |
Illustrative list-price arithmetic that ignores cache-write surcharges. The point is the ratio: routing by task cuts the bill by roughly 4× with no loss where it matters. Across a team of 50 shipping a few features a week, that difference is a budget line.
Tracks by Function
Formats
| Format | Length | Outcome |
|---|---|---|
| Demo session | Half day | One ticket from your backlog taken spec-to-PR live; model and effort choices explained. |
| Team programme | 2 weeks | Workshops per track, context files in your main repos, playbook, templates, skills check, named champions. |
| Programme + measurement | 2 weeks + rollout | Foundations first, then a productivity baseline and pilots via the Claude Code rollout engagement. |
Tech Stack
Related Work
Frequently Asked Questions
What is spec-driven development, in one paragraph?
You write down what you are building and why before any code is generated: a spec with user stories and acceptance criteria, then a technical plan, then a dependency-ordered task list. The AI implements against those files instead of against a one-line chat prompt, and a final check compares the code back to the spec. The result is work that is reviewable, repeatable across engineers, and much less likely to drift. GitHub's open-source Spec Kit and Kiro's built-in specs are the two most common ways teams do it today.
Which tools does the programme cover?
Claude Code and GitHub Copilot in depth, Kiro for teams that use its spec workflow, and Spec Kit on top of any of them. The principles — specs, context files, model and effort choice, review discipline — carry across tools, so a team using a mix of assistants ends up with one shared way of working rather than one per tool.
How do you decide which model and effort level to use?
By task, not by habit. The biggest model at high effort earns its cost on specs, architecture and hard debugging, where a wrong direction wastes days. Well-specified implementation runs well on a mid-tier model at medium effort. Summaries, small edits and ticket triage belong on the fastest model. The page above has the full table for Claude Code and Copilot; in the programme we tune it on your own tasks.
How is this different from vendor training?
Vendor training shows what one tool can do on a demo app. This programme runs on your codebase and your backlog, compares tools honestly, and covers the parts vendors skip: when not to use AI, how to review AI-written code, what it costs per feature, and how to keep engineering IP and customer data out of places it should not go.
How do we know the training worked?
A short skills check before and after, adoption signals from the tools' own admin data, and — for teams that continue into a measurement engagement — workflow metrics such as PR cycle time and test-authoring time against a baseline. The foundations programme is deliberately the step before measurement: it makes sure what you measure afterwards reflects the tools, not a lack of basics.
Can it run remotely, and for how many people?
Yes. It is remote-first with optional on-site days, and it is designed for teams of roughly 10 to 100 engineers, QA and support staff, split into function-specific tracks.
Related Engagements & Reading
What Clients Say
We needed a WhatsApp bot for our clinic chain. Rohit understood the problem immediately and shipped a working solution that our staff could use without training.