Everyone bought the tools. Almost nobody built the practice.

We build a real product with agents every day, test how to do it well the way you would test anything else, and publish the results with the numbers attached. Then we help engineering teams do the same. The experiments that failed are the ones worth reading.

How we work

Everything about building with agents right now is opinion delivered at speed. Ours are hypotheses with a number attached, run on a real codebase, and the number is allowed to say no.

  1. 01HypothesisWe write down what we think a change to how we work will do, and the number that would prove it, before touching anything.
  2. 02ChangeWe make the smallest version that could move that number — a rule, a tool, a way of handing work to an agent — on a real codebase.
  3. 03ResultWe wait the agreed window and read the number as it came out, not as we hoped it would.
  4. 04DecisionKeep it, kill it, or admit we learned nothing. All three get written up and published.

The fourth step is the one most teams skip. An experiment that says "we learned nothing" is a real outcome, and burying it is how a team ends up confidently wrong for a year — which, in a field that turns over this fast, is most of a lifetime.

Latest experiments
The lab

Livnly is the codebase we test all of this on.

A calendar app, in alpha, built almost entirely with agents. It is what keeps the practice honest: every claim we make about working this way has to survive a real product with real users, and the ones that do not survive get published too.

Where Livnly stands
Working together

The same method, pointed at how your team works with agents.

Two days on site. We score how your team works with the tools you already pay for, across six dimensions, and hand back the baseline and a ranked list of what to change. You get the number and the decision, not a deck.