Insights · Spec-driven development

How we connect spec-driven development and AI agents in digital product development

Hypothesis first, then the code

Hypotheses set direction. Specs turn that direction into a verifiable brief. AI agents help us turn it into a working product increment faster.

From fast code generator to productive teammate

AI agents can already write code, create tests, analyse defects, and update documentation in a short time. What they cannot take off our hands is the most important decision in product development: What should we build at all—and why?

The faster software appears, the more important clarity becomes before implementation. AI does not only accelerate good decisions. It also accelerates misunderstandings, wrong assumptions, and unnecessary features.

At Turing 42 we therefore combine hypothesis-driven product development with spec-driven development. Hypotheses set the direction. Specifications translate that direction into a clear, verifiable brief. AI agents help us turn it into a working product increment faster.

The process stays deliberately learning-oriented:

1 Hypothesis
2 Spec
3 Agent loop
4 Product increment
5 Measure
6 Learn

Spec-driven does not mean: back to the requirements bible

Spec-driven development can sound like lengthy requirements documents and long-locked-in specs. That is exactly what we do not mean.

For us, a spec is not an attempt to describe a complete product upfront. It defines the next sensible and verifiable development step precisely enough so people and AI agents can work from the same understanding.

A good spec answers, among other things:

  • Which problem are we solving for which users?
  • Which hypothesis underpins the planned solution?
  • What is in scope for this step—and what is explicitly out?
  • Which behaviour do we expect from the product?
  • Which domain, technical, and regulatory constraints apply?
  • How will we know the implementation is correct?
  • How will we measure whether the underlying hypothesis holds?

Spec depth follows risk. A small UI tweak needs less preparation than a new billing flow, a migration, or a feature that processes personal data.

What matters is not document length, but decision clarity.

Hypothesis and spec have different jobs

Hypothesis-driven product development prevents an idea from instantly becoming a supposedly safe requirement. First we state what we believe—and how we would recognise that the belief is true.

For example:

We believe users of a B2B portal can process relevant cases faster if they no longer have to hunt for information across separate modules. We will know this if the average time to open the right case drops clearly.

The hypothesis describes problem, expected effect, and learning goal. It deliberately leaves open whether the first imagined solution is actually best.

The spec then turns that into an implementable product step. It may describe relevant search objects, permissions, states, interactions, quality requirements, and acceptance criteria. It also records which usage data or feedback we need to check the effect.

The hypothesis says what we want to learn and achieve.
The spec describes what we build and verify next.
The product increment provides evidence for the next decision.

Why specs matter more with AI agents

In classic software development, experienced engineers can spot many ambiguities during implementation, ask questions, and compensate with product or domain knowledge.

AI agents do not automatically share that experience. They work with the context we give them.

An unclear brief therefore often produces code that looks plausible at first glance but misses the real need. A good spec becomes the shared working context for product, UX, engineering, and the agents involved.

It creates guardrails without prescribing every technical decision. It makes desired behaviour testable and stops the agent from silently filling in critical assumptions. At the same time, there is room for technical judgement, sensible alternatives, and new insights during implementation.

For us, spec-driven development is therefore not extra documentation overhead. It is the precondition for embedding AI agents reliably in professional product development.

What this looks like in practice

Back to the B2B portal: the hypothesis was not “we need global search”, but an effect—shorter time to the right case.

The spec for the first step stayed deliberately tight. Excerpt from the brief:

  • Audience: operations staff with access to orders and tickets
  • In scope: cross-cutting search across orders and tickets, including permission filters
  • Out of scope: full-text over attachments, AI ranking, mobile
  • Expected behaviour: a hit opens the case in at most two clicks; empty states and errors are defined
  • Acceptance: automated tests for permissions and hit logic; manual review of the riskiest flows
  • Measurement: time from entry to opening the relevant case in a pilot group

On that basis, agents delivered search, filters, tests, and a docs draft in short loops. People reviewed domain fit, security, and UX—and released what held up.

Measurement showed: time dropped, but less than expected. Interviews explained why—many users did not start with search, but with a favourites list of incomplete cases. The hypothesis was only partly right.

That is the point: without a spec, agents would quickly have built “a search”. With hypothesis, spec, and measurement, we got a learnable outcome—and the next spec targeted the favourites list, not expanding search.

The spec bounds the build brief.
Measurement tests the hypothesis—not only the code.
Refuted assumptions are progress when they arrive early.

Why the agent loop needs specs

In delivery, our senior engineers work with AI agents in iterative loops—often called a Ralph loop: understand, prepare, review, move on. An agent does not get one prompt and return “done”. It plans, implements, runs tests, evaluates results, and improves the solution in bounded passes.

Without a spec, that loop drifts. The agent then optimises what looks locally plausible—not necessarily what should test the hypothesis. With a spec and acceptance criteria, every pass stays aligned to the same guardrails.

How Direction, Loop, and Review work day to day is in our piece on AI agents in the development process. Here the upstream point matters: without a verifiable brief, speed becomes risk.

The human role gets clearer—not smaller

Agents change how work is distributed. They do not take responsibility.

People remain accountable for product judgement, prioritisation, architecture, compliance, and interpreting impact. Engineers orchestrate the process: they create context, make decisions, and safeguard quality—rather than only entering prompts.

The operational split between people and agents belongs in the piece on the AI development process. For spec-driven development, the key point is: the spec is human-owned clarity before speed begins.

More speed—but above all more learning capacity

The biggest advantage of spec-driven development with AI agents is not producing as much code as possible in as little time as possible. More code is not a product goal.

The real gain is shorter, more reliable learning cycles. Assumptions become testable faster, variants more comparable, insights usable earlier—while decisions and quality standards stay traceable.

That is why Turing 42 connects hypothesis-driven product development, senior product engineering, and AI agents:

Not to build arbitrary software faster, but to find out faster which digital product creates real customer value.

AI agents as teammates in the development process

Spec-driven development is part of our AI-first way of working. How agents run from analysis and spec through review and docs—including human-in-the-loop—is in the piece on the AI development process.

Contact & intro call

Write to us or book a slot directly. In an intro call we clarify where you stand, what is blocking you and what makes sense as the next step.

Call: 15-minute video call, pick a slot in the calendar.
Message: Briefly describe your topic, we usually reply within 1–2 business days.

Example topics: new product, modernization, e-commerce, technical due diligence.

Send a message

Compare how we support you again

Book a calendar slot

Free · ~15 min · no sales pressure

Book a 15-min intro call