← All Services

AI Engineering Foundations — Teach Your Team to Actually Use the AI Tools You Already Pay For

A two-week, hands-on programme on your own codebase: spec-driven development from zero, which model and effort level for which task in Claude Code and GitHub Copilot, the folder structure that gives every tool project context, and the token maths behind the bill.

The Problem

Your engineers have Copilot, Claude, Kiro, maybe OpenAI too. Leadership has seen a dozen proofs of concept. Yet nobody can say how much faster the team actually ships, because day-to-day usage is shallow: prompts typed straight into chat with no spec, the most expensive model on every task or the cheapest one on the hard ones, effort left on the default, no repository instructions, so every tool guesses at your conventions. Measuring productivity at this stage only measures the gap in basics. Close the fundamentals first — specs, model and effort choice, context files, review habits — and the gains become real enough to measure.

What You Get

  • ✓Spec-driven development from zero: why specs beat ad-hoc prompting, and the full Spec Kit flow — constitution, specify, clarify, plan, checklist, tasks, analyze, implement, converge — run on a real feature from your backlog
  • ✓Kiro specs side by side with Spec Kit: requirements in EARS notation, design, tasks, and steering files — so teams on either tool follow one process
  • ✓Model and effort playbook for Claude Code and GitHub Copilot: which model and which effort level for spec writing, planning, implementation, debugging, review and support work
  • ✓Repository context set up for your main repos: CLAUDE.md, AGENTS.md, .github/copilot-instructions.md, path-specific instructions and Kiro steering — the single biggest lever on output quality
  • ✓Token and cost basics: context windows, input vs output vs thinking tokens, prompt caching, and a cost model per feature so spend is predictable
  • ✓Separate tracks for Development (C++ and Node/React), QA (specs to test cases) and Support (docs and ticket retrieval)
  • ✓Recorded demo, spec templates and a written playbook your team owns
  • ✓Before-and-after skills check, plus named internal champions who carry the practice after the programme ends

Where This Fits

  1. 1
    Foundations
    This programme. The team learns specs, model and effort choice, context files and review habits.
  2. 2
    Team-wide tool setup, guardrails and a productivity baseline.
  3. 3
    AI systems shipped inside your product by an embedded engineer.

Demo: One Ticket, Spec to Pull Request

A real feature taken from a one-line ticket to a reviewed pull request with Spec Kit — and the model and effort level chosen at each step, with the reason.

Recording in progress — until it is up, the same walkthrough runs live on your own codebase in the free half-day demo.

  1. 0:00Install Spec Kit and run specify init --here in an existing repo
  2. 1:00/speckit-constitution — the project's non-negotiables
  3. 2:00/speckit-specify and /speckit-clarify on a real backlog ticket
  4. 3:30/speckit-plan with Opus at high effort, /speckit-tasks
  5. 5:00/speckit-implement on Sonnet at medium effort
  6. 6:30/speckit-converge, review the diff, open the PR

Spec-Driven Development, From Zero

Most teams use AI coding tools by typing a one-line request into chat. That works for a small edit and falls apart for a feature: the agent guesses the requirements, picks its own architecture, and nobody can review the result against anything. Spec-driven development fixes the order of work — what and why first, then how, then tasks, then code — and keeps each step in a file that humans review and the agent implements against.

The Spec Kit flow

Install once: uv tool install specify-cli, then, inside an existing repo, specify init --here --integration claude (or copilot). The steps run inside the agent's chat, not the terminal. Copilot and Claude Code both invoke them as /speckit-specify and so on; Spec Kit's docs write the same steps as /speckit.specify.

StepWhenWhat it produces
/speckit-constitutionOnce per projectPrinciples every later step is checked against: testing rules, security rules, architecture constraints.
/speckit-specifyEvery featureThe what and the why: user stories and acceptance criteria. No tech stack here.
/speckit-clarifyQuality gateThe agent asks up to five targeted questions and writes the answers back into spec.md.
/speckit-planEvery featureThe how: stack, architecture and constraints become plan.md and supporting design files.
/speckit-checklistQuality gate“Unit tests for requirements” — checks the spec is complete and unambiguous before work is split.
/speckit-tasksEvery featureA dependency-ordered tasks.md: setup, foundations, one phase per user story, polish.
/speckit-analyzeQuality gateRead-only consistency check across spec, plan and tasks before any code is written.
/speckit-implementEvery featureExecutes tasks.md phase by phase, respecting dependencies and parallel markers.
/speckit-convergeEvery featureCompares the code with the spec; appends missing tasks. Repeat implement → converge until Converged.

Small features take the short path: specify → plan → tasks → implement → converge. Production features add the three quality gates (clarify, checklist, analyze).

Kiro specs, for teams on Kiro

Kiro builds the same idea into the IDE. Each spec lives in .kiro/specs/<feature>/ as requirements.md (or bugfix.md), design.md and tasks.md. Acceptance criteria use EARS notation, which keeps them testable:

WHEN an auditor exports a report for a closed period
THE SYSTEM SHALL produce a PDF signed with the firm's certificate
AND record the export in the audit log with user, time and checksum.

What a good spec contains

Weak specStrong spec
“Add export to reports.”User story, who needs it and why, and 3–6 acceptance criteria a tester can check.
Mixes UI details, database choice and goals in one paragraph.Spec says what and why; the stack and architecture go in the plan.
Silent on errors, permissions and limits.Names edge cases: empty data, no permission, large files, time-outs.
No link to existing rules.Inherits the constitution: test coverage, security and performance rules.

The Folder Structure That Gives Every Tool Context

Output quality depends less on the model than on what the model knows about your repo. These files are how you tell it — once, in version control, for every engineer and every tool.

your-repo/
├── AGENTS.md                         ← shared instructions every agent reads
├── CLAUDE.md                         ← Claude Code project memory (Copilot reads it too)
├── .specify/                         ← Spec Kit
│   ├── memory/constitution.md        ← project principles, written once
│   ├── templates/  scripts/
│   └── feature.json                  ← which feature is active
├── specs/
│   └── 001-export-audit-report/      ← one folder per feature
│       ├── spec.md                   ← what & why
│       ├── plan.md                   ← how (+ research, data model, contracts)
│       ├── tasks.md                  ← ordered work items
│       └── checklists/requirements.md
├── .claude/                          ← Claude Code
│   ├── skills/speckit-*/SKILL.md     ← Spec Kit skills (specify init --integration claude)
│   ├── agents/  commands/
│   └── settings.json                 ← permissions, hooks, per-model effort
├── .github/                          ← GitHub Copilot
│   ├── copilot-instructions.md       ← repo-wide instructions
│   ├── instructions/cpp.instructions.md   ← path-specific, applyTo: "src/**/*.cpp"
│   └── skills/speckit-*/SKILL.md     ← Spec Kit skills (specify init --integration copilot)
└── .kiro/                            ← Kiro
    ├── specs/export-audit-report/
    │   ├── requirements.md           ← user stories + EARS acceptance criteria
    │   ├── design.md
    │   └── tasks.md
    └── steering/                     ← Kiro's equivalent of instruction files

Copilot reads .github/copilot-instructions.md, path-specific .github/instructions/*.instructions.md files, the nearest AGENTS.md, and a root CLAUDE.md. Keep shared rules in AGENTS.md so a team using several tools maintains one source of truth.

Which Model and Which Effort, for Which Task

Two dials matter. The model sets the ceiling on capability and the price per token. The effort level sets how many tokens the model spends thinking, calling tools and writing — on Anthropic's current models, effort is the main control for depth and cost. The expensive mistake is running everything at one setting.

Claude Code

Switch model with /model opus, /model sonnet, /model haiku or /model fable; set effort with /effort high or claude --effort xhigh. The opusplan alias plans with Opus and implements with Sonnet — a good default for spec-driven work. Per-model effort can be pinned in .claude/settings.json under modelSettings, and a skill or subagent can set its own effort in frontmatter.

TaskModelEffortWhy
Constitution, spec, clarifyOpus 5.5highWrong requirements waste days; this is where depth pays.
Plan and task breakdownOpus 5.5 (or opusplan)highArchitecture choices and task ordering need reasoning.
Implement a well-specified taskSonnet 5.5mediumThe spec already carries the thinking; speed matters more.
Implement a hard or long taskSonnet 5.5 or Opus 5.5highRaise effort before switching model.
Hard C++ bug, legacy module, cross-cutting refactorFable 5.1 or Opus 5.5high → xhighLong reasoning chains across unfamiliar code.
Autonomous runs over ~30 minutesOpus 5.5 or Fable 5.1xhighAnthropic's documented use case for xhigh.
Code review of AI-written changesOpus 5.5highA reviewer should be at least as capable as the author.
Small edits, summaries, ticket triage, subagentsHaiku 4.5 or Sonnet 5.5— / lowFastest and cheapest; Haiku has no effort setting.

GitHub Copilot

Choose the model in the chat input's model picker; for reasoning models, the arrow next to the model name opens a Thinking Effort menu (None, Low, Medium, High). Auto lets Copilot route by task complexity — fine for everyday chat, but pin the model for spec and review work so results are repeatable. The same Claude models are available in Copilot alongside OpenAI's GPT and Codex models, Gemini and others; which ones you see depends on your plan and your organisation's policy.

TaskModel in CopilotThinking effort
Spec, plan, architecture questionsClaude Opus 5.5 or a top GPT modelHigh
Agent-mode implementation from tasks.mdClaude Sonnet 5.5 or a Codex modelMedium
Inline completions and quick chatAuto or a fast model (Haiku 4.5, a mini model)Low / None
Reviewing a pull requestClaude Opus 5.5 or equivalentHigh

Larger models use more of a Copilot plan's premium allowance, so the routing above is also a budget decision.

Effort levels explained

LevelUse it forTrade-off
lowSimple, well-defined tasks; subagents; chat.Fastest and cheapest; may under-think hard problems.
mediumWell-specified agentic coding. Default on Opus 5.5 and in Claude Code on Sonnet 5.5.The balance point for most implementation.
highComplex reasoning, difficult coding, specs and plans. API default on most models.Spends what the task needs.
xhighLong-running agentic work, roughly 30+ minutes.Meaningfully more tokens than high.
maxFrontier problems only.Often large cost for small gains; can overthink.

Rule of thumb: raise effort before switching to a bigger model, and lower it only where you have checked that quality holds.

Tokens and Cost, the Basics

  • Token: the unit models read and bill in. On Claude's current tokenizer, 1M tokens is roughly 555,000 words.
  • Context window: how much the model can hold at once — 1M tokens on Fable 5.1, Opus 5.5 and Sonnet 5.5; 200K on Haiku 4.5. Bigger is not free: every token in context is billed on every turn.
  • Input vs output: output costs five times input on current Claude models, and thinking counts as output. That is why effort level moves the bill.
  • Prompt caching: repeated context (instruction files, the spec, unchanged code) is read from cache at a fraction of the input price — 10% on most models, 5% on Opus 5.5, 2.5% on Fable 5.1.
  • Batch: non-urgent jobs through the Batch API are 50% off.

Current Claude models

ModelBest forInput / output per 1M tokensContextDefault effort
Claude Fable 5.1Hardest reasoning, long agentic runs$10 / $501Mhigh
Claude Opus 5.5Specs, plans, agentic coding, review$4 / $201Mmedium
Claude Sonnet 5.5Everyday implementation$2 / $101Mhigh (medium in Claude Code)
Claude Haiku 4.5Fast, cheap, high-volume$1 / $5200Knot supported

API list prices, last verified 29 September 2026 against Anthropic's models overview. Seat-based plans (Claude Team and Enterprise, Copilot) bill differently; the relative costs still hold.

Worked example: one feature, two ways

A mid-sized feature: about 150K input and 40K output tokens for spec, plan and tasks, then about 2M input (80% cached) and 150K output for implementation.

ApproachSpec + planImplementationTotal
Everything on Fable 5.1$3.50$11.90≈ $15.40
Opus 5.5 plans, Sonnet 5.5 implements$1.40$2.62≈ $4.02

Illustrative list-price arithmetic that ignores cache-write surcharges. The point is the ratio: routing by task cuts the bill by roughly 4× with no loss where it matters. Across a team of 50 shipping a few features a week, that difference is a budget line.

Tracks by Function

Development
Spec Kit on real tickets; C++ track for legacy-code comprehension, test generation and build-failure triage; Node/React track for agentic implementation and AI-assisted review.
QA
Acceptance criteria to test cases, test automation authoring, defect triage and de-duplication — with specs as the single source of truth.
Support
Answers grounded in product docs and past tickets, ticket summaries, log analysis, and cleaner escalations to engineering.

Formats

FormatLengthOutcome
Demo sessionHalf dayOne ticket from your backlog taken spec-to-PR live; model and effort choices explained.
Team programme2 weeksWorkshops per track, context files in your main repos, playbook, templates, skills check, named champions.
Programme + measurement2 weeks + rolloutFoundations first, then a productivity baseline and pilots via the Claude Code rollout engagement.

Tech Stack

Spec KitKiroClaude CodeGitHub CopilotClaude Opus 5.5Claude Sonnet 5.5Claude Fable 5.1Claude Haiku 4.5AGENTS.mdMCP
Timeline
Half-day demo, or a 2-week team programme

Related Work

Agentic OS — Force-Directed Map of a 387-Skill Claude Code Setup
An agent setup grows one plugin at a time until it has hundreds of skills, and nobody can see what it actually contains.
claude-autodev — Autonomous 8-Stage Dev Pipeline for Claude Code
Agentic coding tools stop at a diff.

Frequently Asked Questions

What is spec-driven development, in one paragraph?

You write down what you are building and why before any code is generated: a spec with user stories and acceptance criteria, then a technical plan, then a dependency-ordered task list. The AI implements against those files instead of against a one-line chat prompt, and a final check compares the code back to the spec. The result is work that is reviewable, repeatable across engineers, and much less likely to drift. GitHub's open-source Spec Kit and Kiro's built-in specs are the two most common ways teams do it today.

Which tools does the programme cover?

Claude Code and GitHub Copilot in depth, Kiro for teams that use its spec workflow, and Spec Kit on top of any of them. The principles — specs, context files, model and effort choice, review discipline — carry across tools, so a team using a mix of assistants ends up with one shared way of working rather than one per tool.

How do you decide which model and effort level to use?

By task, not by habit. The biggest model at high effort earns its cost on specs, architecture and hard debugging, where a wrong direction wastes days. Well-specified implementation runs well on a mid-tier model at medium effort. Summaries, small edits and ticket triage belong on the fastest model. The page above has the full table for Claude Code and Copilot; in the programme we tune it on your own tasks.

How is this different from vendor training?

Vendor training shows what one tool can do on a demo app. This programme runs on your codebase and your backlog, compares tools honestly, and covers the parts vendors skip: when not to use AI, how to review AI-written code, what it costs per feature, and how to keep engineering IP and customer data out of places it should not go.

How do we know the training worked?

A short skills check before and after, adoption signals from the tools' own admin data, and — for teams that continue into a measurement engagement — workflow metrics such as PR cycle time and test-authoring time against a baseline. The foundations programme is deliberately the step before measurement: it makes sure what you measure afterwards reflects the tools, not a lack of basics.

Can it run remotely, and for how many people?

Yes. It is remote-first with optional on-site days, and it is designed for teams of roughly 10 to 100 engineers, QA and support staff, split into function-specific tracks.

Related Engagements & Reading

Claude Code Consultant →
The next step after foundations: team-wide Claude Code setup, guardrails and measurement.
Fractional AI Engineer →
When you want AI systems built inside your product, not just a faster team.
Claude Code plugins and context engineering →
How skills, plugins and context budgets fit together in a real setup.
Cutting LLM token costs →
Context compression techniques that lower the bill without lowering quality.
Testimonials

What Clients Say

Rohit delivered our MVP in 5 weeks — on budget and ahead of schedule. His architecture decisions saved us from rewriting everything when we scaled.

Arjun Kapoor
Founder, NovaByte Labs
MVP Development

We needed a WhatsApp bot for our clinic chain. Rohit understood the problem immediately and shipped a working solution that our staff could use without training.

Priya Mehta
CTO, MediConnect Health
WhatsApp Bot
Book a Free Half-Day Demo