Axioms Lab

Axioms Lab
Idle

Autopilot

Replays a fixture through the real pipeline, judges the board, and proposes a better version of the pipeline algorithm. Nothing is applied until you say so.

Sources

?
to

Run

?
? · min ?

Models

Yours to choose. The autopilot never touches these. ?

The fact-check judge is a re-check: the same claims checked again, twice, by this model. Empty makes it a consistency check rather than a second opinion.

Judging

How much each axis counts ?

The goal

What a perfect board is. Quoted verbatim into the judge's and the proposer's prompts, so editing it re-aims the whole loop.

Lineage

The lineage draws itself here once a run starts: one column per generation, the rival changes inside it, and every replay's own progress.

Record

Every run so far ?

Nothing has run yet.

Live log

Pipeline

Drag to pan · wheel to zoom · click a step

Replay

Streams a fixture's audio through a real relay as a real debate, at wall-clock speed, so you can watch and hear the board fill.

Runs on your machine. A replay streams the fixture's audio lanes to a relay and plays them out loud, and this runner holds no audio. Start the Lab locally for this.

Replays here

?

Every debate in the environment selected below, newest first, with what it has cost so far. The arrow opens it in that environment's app.

Fixture

Where it runs

?

Production is not offered: a replay there is a real debate on the public board, billed for the length of the fixture.

Set it up like a debate

?

Every field is optional. Empty means the fixture's own title and question, and the column defaults the wizard starts from.

Ingest & label

A YouTube debate in, a labelled fixture out: the audio, who spoke when, and the names you put to each voice.

Runs on your machine. Ingest downloads a recording and keeps gigabytes of working copies, none of which belongs on a server. Start the Lab locally for this. Every fixture you save there reaches the deployed autopilot on its own.

New fixture

?

Videos on this machine

?

Each video you ingest leaves a working copy here: the mixed audio and Deepgram's segments. Labelling one cuts a fixture from it, which the Autopilot and Replay tabs use.

Sources and the hold-out

Sources are where a change is found. The hold-out is where it is believed.

A change is picked because it scored well on the sources, which is grading its own homework: it may have learnt that quarter hour rather than got better at debates. So a winner is re-measured on the hold-out, minutes the search never saw. Losing there means overfit, and it is thrown out.

It must not overlap a source. Times are points in the recording. A fixture marked bucket was uploaded from a laptop and holds its transcript only; one marked disk is here whole, audio included.

Generations, shots and candidates

One generation: run every source, score the board, keep or revert, propose the next change. It takes about as long as the slice, because the replay is real time.

Candidates are rival changes tried in the same generation; the best survives and becomes the next generation's parent. Shots are repeat readings; a score is the mean of its shots less their spread.

The run starts from the pipeline algorithm staging runs right now. The cap covers the whole run. At the end, a run that beat its start proposes a new version and asks; it never applies one.

The fact checker is measured only when its step is unlocked on the Pipeline page: two checks a shot, then none. Money for a number nobody can act on is money wasted.

How the score is made

Each shot's positions score is 0.65 the judge's axes under these weights, plus 0.35 the deterministic checks that cost nothing. The fact-check and viewer groups have their own judges and their own axes.

A weight of 0 silences an axis; 2 doubles it. The default weights put attribution and position fidelity first, because a wrong mouth on a public surface is the failure the product cannot afford.

The goal document is the rubric all three judges read. It is versioned content: what ships is in the repo, and an edit here is kept in the database and overrides it until you go back.

Which Claude

These pick who is billing, not which model. Claude Code is the claude binary on a signed-in machine, under the subscription already paid for; it exists on your laptop and not on the deployed runner. OpenRouter is the paid API the product itself uses, metered per token, and Claude through OpenRouter is a metered call, not your subscription.

The pipeline starts on OpenRouter because that is what live debates run on. Production always runs on OpenRouter whatever is chosen here.

Reading the record

Generations run left to right, so the line across a row is the improvement. Only the settings change between them.

Top rows are each recording's score out of 10, coloured by band: under 5, 5 to 7, 7 and above. Aggregate is their weighted mean, Change the move from the generation before.

Kept became the next parent, reverted was put back. A difference smaller than the shots' own standard error is a tie and proves nothing.

Where a replay runs

Local streams to the relay on this machine, ws://localhost:8787/ws. Staging streams to the deployed staging relay, so the debate appears on the staging site exactly as a real one would. Same database as local.

Production is absent on purpose. A replay mints a genuine debate on the public board and pays to transcribe and analyse the whole fixture.

Setting a replay up

The same framing and model choices the debate wizard asks for, minus the audio step and the speakers, which the fixture supplies. Leave a field empty and the fixture's manifest, or the column default, fills it.

What a replay does

Streams a fixture through the relay as a real debate: real sockets, transcription and analysis, at wall-clock speed. Nothing is graded and nothing is applied. Nothing reaches a paid provider until you press Go live in the app; after that it costs what a real debate of that length costs.

To end one, end or delete the debate in the app; the Lab notices and stops the replay.

What ingest does

Downloads the audio, asks Deepgram who spoke when, plays you a sample of each voice, and you put a name to each one. What comes out is a set of isolated per-speaker lanes, which is what the relay expects from real speakers on real microphones, plus a transcript the autopilot grades against. Saving uploads the transcript and the manifest to the Lab's bucket; the audio stays here.

What is kept here

One working copy per video you ingest: the mixed audio and Deepgram's segments. They stay on this machine, and are never uploaded, committed or deployed. Labelling one cuts a fixture from it. They are large; deleting a video frees the disk and loses nothing but the ingest time.