Shashi Techlogues

AI + SDD + TDD: Make the AI Prove It Works Before It Codes

AI coding agents have changed the economics of software development.

A feature that once took a developer several hours to implement can now be generated in minutes. An AI agent can create files, modify existing code, write database queries, refactor modules, and even run the test suite.

But there is a problem that becomes more important as AI gets better at writing code:

How do we know the code is actually correct?

The obvious answer is tests.

But there is a subtle problem with the usual AI workflow:

Requirement
    ↓
AI writes implementation
    ↓
AI writes tests
    ↓
Tests pass

The same system that wrote the implementation also gets to decide how the implementation should be tested.

That is backwards.

A better workflow is:


The important idea is simple:

Don't let AI define correctness after it has written the code. Define correctness first, then make AI satisfy it.

This is where Spec-Driven Development (SDD) and Test-Driven Development (TDD) become particularly interesting when combined with AI.

The Core Idea

The workflow I have been experimenting with is:

Specification → Tests → Implementation → Verification

Each stage has a different purpose.

Layer Question
Specification What should the system do?
Tests What does "correct" mean in executable form?
Implementation How should we make it happen?
Verification Does the implementation still satisfy the contract?

AI can participate heavily in all four stages.

But it should not have unrestricted authority over all four.

The human should remain responsible for intent.

The tests should establish the contract.

The AI should have freedom over the implementation.

And automated tooling should perform much of the verification.

SDD: Start With Intent

Spec-Driven Development starts by describing what the system is supposed to do before jumping into implementation.

Suppose we need a payment API.

The requirement might be:

Create a payment and return a payment ID. Repeating the same request with the same idempotency key must not create another payment.

A technical specification might make that more explicit:

POST /payments

Given a valid payment request:
    create a payment
    return HTTP 201
    return the payment ID

Given the same idempotency key:
    return the existing payment
    do not create another payment

Given an invalid request:
    return HTTP 400

Given a payment-provider failure:
    return an appropriate error
    do not leave the system in an inconsistent state

This is not implementation.

It doesn't say:

Use Redis.
Use PostgreSQL.
Create a PaymentService class.
Use distributed locks.

Those are implementation decisions.

The specification describes behavior and constraints.

That distinction becomes extremely important when an AI agent is doing the implementation.

TDD: Turn the Specification Into an Executable Contract

Now we ask AI to write the tests.

Not the production code.

A prompt could be:

Here is the specification.

Create a comprehensive test suite for this behavior.

Cover:
- happy paths
- boundary conditions
- invalid inputs
- failure modes
- idempotency
- concurrency where relevant

Prefer behavior-level assertions over implementation details.

Do not implement the feature yet.

The AI might produce:

it("creates a payment for a valid request", async () => {
  const response = await createPayment(request);

  expect(response.status).toBe(201);
  expect(response.body.paymentId).toBeDefined();
});

it("does not create a duplicate payment for the same idempotency key", async () => {
  const first = await createPayment(request);
  const second = await createPayment(request);

  expect(second.body.paymentId)
    .toBe(first.body.paymentId);

  expect(await countPayments()).toBe(1);
});

it("rejects an invalid payment request", async () => {
  const response = await createPayment(invalidRequest);

  expect(response.status).toBe(400);
});

Now something important happens.

We stop and review the tests.

We don't immediately ask AI to implement the feature.

The Human Review Is the Critical Step

AI-generated tests aren't automatically correct.

AI can misunderstand requirements just as easily as it can misunderstand implementation.

So this is where human judgment belongs:

Specification
      ↓
AI generates tests
      ↓
Human asks:
      │
      ├── Does this represent the requirement?
      ├── Are important cases missing?
      ├── Are these assertions behavioral?
      ├── Did AI invent an assumption?
      └── Is anything ambiguous?

This is a much better use of human engineering time than manually inspecting hundreds of lines of AI-generated boilerplate.

The human is reviewing what the software should do, rather than initially reviewing how the software was written.

Now Run the Tests

Once the tests are approved, run them.

They should fail.

That's a feature, not a problem.

FAIL

ReferenceError:
calculatePrice is not defined

We now have something valuable:

An executable specification that the current system does not satisfy.

This is classic TDD:

RED
 ↓
GREEN
 ↓
REFACTOR

But AI changes the economics of the loop.

The AI can now take the failing tests and implement the functionality.

Give the AI a Much Smaller Problem

The next prompt becomes:

The tests have been reviewed and approved.

Implement the minimum production code required to make the existing tests pass.

Rules:

- Do not modify the tests.
- Do not weaken assertions.
- Do not remove tests.
- Do not skip tests.
- Do not change unrelated behavior.

Run the tests after implementation and report the result.

This is a very different problem from:

"Build this feature."

The AI is no longer being asked to simultaneously determine:

  • what the requirement means,
  • what the expected behavior is,
  • how it should be tested,
  • and how it should be implemented.

Those responsibilities have been separated.

The AI gets a contract.

Its job is to satisfy it.

The Workflow

AI-assisted SDD + TDD workflow

Requirement
    ↓
Technical Specification
    ↓
AI writes Tests
    ↓
Human reviews Contract
    ↓
Run Tests → FAIL
    ↓
AI writes Implementation
    ↓
Run Tests
    ↓
PASS

And if they fail:

FAIL
 ↓
Investigate
 ↓
Fix implementation
 ↓
Run tests again

Not:

FAIL
 ↓
Change the test
 ↓
PASS

That second workflow destroys the whole point.

Never Let the AI Change the Contract to Make the Code Green

This is probably the most important rule in the entire approach.

Suppose the approved test says:

expect(calculatePrice(1000, "premium")).toBe(800);

The implementation returns:

850

The test fails.

A bad AI workflow is:

Test fails
    ↓
AI changes test to expect 850
    ↓
Test passes

Now the test is merely validating the implementation.

The correct workflow is:

Test fails
    ↓
Investigate implementation
    ↓
Fix implementation
    ↓
Test passes

Once approved, the tests should be treated as part of the specification.

The implementation adapts to the contract, not the other way around.

How SDD and TDD Fit Together

This is where I think the combination gets particularly interesting.

SDD answers:

What should the system do?

TDD answers:

How can we express that behavior as an executable contract?

AI answers:

How can we efficiently implement that contract?

CI answers:

Does the implementation still satisfy it?

SDD, TDD, AI and verification form different layers of the workflow.

This gives every layer a clear responsibility.


Why This Matters More With AI

You might ask:

"If AI can already write tests, why bother doing this?"

Because AI makes implementation cheap.

And when implementation becomes cheap, correctness becomes the bottleneck.

Imagine an AI agent can generate 1,000 lines of code in a few minutes.

The question is no longer:

"Can we write this code?"

The question becomes:

"How do we know this is the code we actually wanted?"

That is fundamentally a specification problem.

AI Is Very Good at Finding Edge Cases

There is another advantage to putting AI on the test-generation side.

AI is surprisingly useful at generating possible edge cases.

Normal input
Empty input
Null input
Minimum value
Maximum value
Duplicate request
Concurrent request
Timeout
Retry
Partial failure
Malformed response
Dependency unavailable
Unexpected dependency response

A developer can ask:

"What cases should we consider?"

AI can generate a large list quickly.

But there is an important distinction:

AI generates possibilities. The developer decides which possibilities are part of the contract.

That is a much healthier division of responsibility.

Tests Should Describe Behavior, Not Implementation

There is another trap.

If the AI is generating tests, it can easily generate overly implementation-specific tests.

For example:

expect(service.cache.size).toBe(1);

This test says something about the internal implementation.

It effectively says:

"You must implement caching this particular way."

A better test is:

expect(await service.get("user-123"))
  .toEqual(expectedUser);

This says:

"The system must return the correct user."

The AI remains free to decide whether the implementation uses Redis, an in-memory cache, a database query, or some completely different mechanism.

Good tests constrain behavior, not implementation.

This is particularly important when AI is responsible for the implementation.

The AI Development Loop Changes

Traditional software development often looks like:

Requirement
    ↓
Developer
    ↓
Design
    ↓
Code
    ↓
Tests
    ↓
Review

AI-assisted development can become:

Requirement
    ↓
Specification
    ↓
AI generates tests
    ↓
Human approves contract
    ↓
AI generates implementation
    ↓
Automated verification
    ↓
Human reviews important decisions

The human moves upward in the abstraction stack.

Instead of spending most of their time typing implementation code, they spend more time defining:

  • behavior
  • constraints
  • architecture
  • invariants
  • acceptance criteria
  • trade-offs

That's a meaningful change in the role of the software engineer.

A Better Way to Think About AI Coding Agents

An AI coding agent should not be thought of simply as:

"A really fast developer."

A better mental model is:

An implementation engine operating inside a set of constraints.

Those constraints can come from:

Architecture
    +
Specification
    +
Tests
    +
Type system
    +
Linting
    +
CI

The more of these constraints are explicit and executable, the safer it becomes to give the AI autonomy.

From TDD to Spec-Driven AI Development

For larger projects, I would extend the workflow further:

Product Requirement
        ↓
System Specification
        ↓
Architecture / Design
        ↓
Acceptance Criteria
        ↓
Integration Tests
        ↓
Unit Tests
        ↓
Implementation
        ↓
Verification
        ↓
Review
        ↓
Merge

AI can assist at almost every step.

But the level of autonomy should be different.

Humans should own

  • Product intent
  • Business rules
  • Architectural boundaries
  • Security requirements
  • Important trade-offs
  • Acceptance criteria

AI can assist with

  • Specification drafting
  • Test generation
  • Implementation
  • Refactoring
  • Documentation
  • Edge-case discovery
  • Debugging

Automated systems should enforce

  • Compilation
  • Unit tests
  • Integration tests
  • Static analysis
  • Security checks
  • Build validation
  • Deployment gates

This gives us a much more robust division of responsibility.

The Specification Becomes the Interface to the AI

There is a deeper implication here.

Today, the primary interface between a developer and a codebase is the code itself.

With AI agents, that interface can increasingly become:

Specification
      +
Architecture
      +
Tests
      +
Repository

Instead of telling the AI exactly how to implement something:

"Create this class, inject this dependency, call this method, update this repository..."

we can increasingly say:

"This is the behavior the system must provide."

The AI figures out the implementation.

That is a significant shift.

Traditional AI Coding vs Spec-First AI Coding

The key change: the executable contract is established before implementation.

The difference can be summarized simply:

Traditional AI workflow Spec-first AI workflow
AI writes code first AI writes tests first
Tests may reflect implementation Tests define expected behavior
Correctness is discovered later Correctness is defined earlier
Human reviews implementation heavily Human reviews the contract first
AI has more freedom to interpret requirements AI operates within explicit constraints



The Most Important Rule

If I had to reduce this entire approach to one rule, it would be:

Never ask AI to implement a non-trivial requirement until there is an executable definition of what "correct" means.

And ideally:

Let AI write the first version of that definition, but make a human approve it before AI writes the implementation.

The complete model becomes:

Human defines intent
        ↓
SDD captures the intent
        ↓
AI translates it into tests
        ↓
Human approves the contract
        ↓
TDD establishes RED
        ↓
AI implements
        ↓
Tests establish GREEN
        ↓
CI continuously verifies

The Bigger Shift

The interesting future of AI-assisted software engineering isn't simply:

"AI writes code faster."

It is:

"AI can operate inside executable specifications."

That's a much more powerful idea.

We don't need to give AI unrestricted freedom and hope that it produces the right result.

We can give it a clearly defined contract and let it explore the implementation space inside that contract.

The human defines the destination.

The specification defines the boundaries.

The tests define correctness.

The AI navigates the implementation.

And the automated toolchain checks whether it actually arrived.

Don't make the AI write code and then prove that it works.

Make it define the behavior, prove that the definition fails, and only then let it write the code.

That is where SDD + TDD + AI starts to look less like AI-assisted coding and more like a new software-development workflow.


What do you think? Is this the direction AI-assisted software development should take, or does the additional specification and test layer introduce too much process? I'd be particularly interested in hearing how others are using coding agents with TDD or Spec-Driven Development.