LA FORJAEnter the studio

An adversarial learning studio · high-school & college mathematics

Getting the right answer
is not enough.

Here you don’t answer math problems — you build one, watch AI reviewers attack it with evidence, repair it, defend it in writing, and publish it with a passport that records the whole fight.

The studio is the workshop itself — a broken problem is already waiting on the bench. No account, no setup, five minutes: load it, find the flaw, repair it, and walk out with a passport.

The demo challenge — one stem, two defensible answers

A family has two children. It is known that one of them is a boy. What is the probability that both children are boys?

Reading A

“At least one of the two is a boy”

1/3
Reading B

“One specific child is a boy”

1/2

Two readings, two answers: the item is broken. You repair it.

Why a forge?

When AI arrived in classrooms, something wonderful happened: students stopped getting stuck. The question that used to block you until the next class now gets unblocked in thirty seconds. But something quieter happened too. The same shortcut that unblocks you also lets you step around the problems that demand real effort — the slow maturation of a concept, the uncomfortable stretch where abstraction is built. You reach the answer without ever owning the reasoning.

If students use AI as an answer machine, they will not learn. The tool is not the problem. The direction of the interaction is.

And there is a harder question underneath, one every student deserves to hear: you are not graduating into the world you were prepared for. You are graduating into a world where intelligence is suddenly everywhere. So who becomes more valuable when everyone has access to intelligence? The person who can build — and defend what they built under pressure.

That is why this studio is a forge and not a tutor. Here the AI never answers for you: you author the problem, the AI attacks it with evidence, and the repair and the defense are yours alone. Writing one good question demands deeper understanding than answering twenty — you must master the content, anticipate how others go wrong, and design every wrong option around a mistake a real student would make. Builder thinking, applied to mathematics.

Read the whole story →

One item, six stations

Publication is earned, never granted. Follow one problem through the forge, top to bottom — each station leaves a record the passport keeps.

  1. FORGEalways on

    Start with a broken problem

    You load a team-authored item with a deliberate defect hidden in it. Repair-first: your first move is never an empty form.

  2. GAUNTLETlive with an API key

    Send it into the gauntlet

    Three AI reviewers and a deterministic probe attack the item at the same time, each bound to an evidence contract. A claim without evidence never counts.

    the deterministic probe runs without a key

  3. FRACTURElive with an API key

    Face the counterexample

    The strongest finding is not an opinion — it is two honest readings of your stem that force two different answers, shown so you can re-execute them yourself.

  4. HAMMERalways on

    Repair it — version 2

    You rewrite the stem so only one reading survives. A repair never overwrites: version 1 stays on record, and the full check history re-runs against version 2.

    re-judging semantic checks needs the key

  5. PROOFlive with an API key

    Defend it in writing

    Fixing it is not enough — you show you understand why it was broken. Two written questions, scored on a three-part rubric with quoted evidence for every score.

  6. STAMPalways on

    Publish with a passport

    The surviving item ships with its diploma: every attack, every re-run, the verdicts, the rubric, every version. Auditable by anyone.

The AI does not generate the initial item and does not hand over a canonical solution to copy; it returns challenges and evidence — the repair and the defense belong to the student.

Three kinds of checks, three different promises

The studio never claims more than a check can keep. Each class carries its own promise, and the difference stays visible everywhere.

deterministic

Cannot regress

Schema invariants, answer-count checks, reproducible solver runs. Strict non-regression: version 2 cannot reintroduce the failure.

counterexample

Re-executed, and it blocks

A concrete construction — two readings, two answers. It is re-run on every new version; while it still holds, the item does not publish.

semantic

Re-judged, never absolute

Plausibility judgments are re-adjudicated on every version and shown in the passport. They are never described as a guarantee.

Every accepted check is re-run on each new version. The system guarantees execution of the history and non-regression of deterministic invariants; semantic judgments are re-adjudicated and remain visible in the passport.

The only guarantee this studio makes

Stated plainly

The whole pipeline — three reviewers, adjudication, written defense, passport — runs end to end against live GPT-5.6, and the evaluation has run for real. On the 14-item holdout the single-reviewer baseline finds all 12 planted defects but flags both clean items; the full gauntlet finds 6–7 with at most one false positive. Exact counts, committed to the repo — no number anywhere came from a run that did not happen.

Read the numbers

The rules of the house

The forge keeps a few promises it will not trade away. Nobody is on file: authors appear as random pseudonyms, and no name, email, school, city or age field exists in any schema, form or column — not optional, absent. Everything is ours to give: every item is a team-authored original released under CC-BY, with no third-party exam content of any kind. No exam owns the mechanism: it was designed against the constraints of a real high-stakes exam — that is where the assessment expertise comes from — and built to be exam-agnostic; today’s arenas are probability, statistics and geometry. Your text is handled like evidence: item text is delimited in every prompt, reviewers get no tools and no open network, and input is capped and rate-limited.

“Defending a solution in front of peers who genuinely wanted to find the flaw remains one of the richest learning experiences of my life. That forum is gone. LA FORJA is our attempt to relight that flame.”

— from the founder’s story
About the project

Start with a broken problem.

Onboarding is repair-first: your first move is loading a defective item and watching it fracture — not staring at an empty form.

Enter the studio