Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SeaWolf-AI 
posted an update 5 days ago
Post
3391
The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.

Three boards get the same piece order, and on every move the same proposal and the same noise — a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.

The text-writing gates get through 15–22 moves. The generation-free gate gets through 40–50. The boards that stop simply run out of clock.

It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock — 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration — a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism — a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.

The curve in the ZTC panel is real online fitting, scored prequentially — predict first, learn after — with base weights untouched. Not recursive self-improvement.

Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.

KO / EN / ZH.

FINAL-Bench/Tetris-JEV-LAYA-ZTC
FINAL-Bench/ZTC-Judge-27B

The Laya board is a threshold result, not a calibration result.

In index.html, THR = probit(0.5 + ECE). For Laya that is probit(0.9985) = 2.97, against a separation of sqrt(2)*probit(0.735) = 0.89.

So Laya re-rolls 99.4% of good moves and 99.97% of bad ones. Its AUC barely enters the board. The 14.3 moves per 60 s falls straight out of 2.1 s x (1 + 0.997).

ECE is unsigned, so the page has to pick a direction, and you flag that. The bigger choice is cutting on the raw score at all. Anyone running a scorer with AUC 0.735 and ECE 0.5 sets the cut on a held-out slice, at a quantile. Monotone miscalibration costs nothing there.

Laya and JEV share SEP, latency, the draw and the rng seed (mulberry(seed+7) on both). Set THR_laya = THR_jev and the two boards play the same game, move for move. The calibration axis goes to zero.

The separation that survives is the clock: about 3.22 s vs 0.094 s per move once each gate's re-roll rate is in.

One question on the ZTC curve. learner.step(mv.f, good?1:0) trains on good, the simulator's own top-15% pool label. A deployed gate never sees that label.

What would the per-move label be outside the sim, line clears over the next N moves?

·

Sharp read, thanks for going through the source.

Both points are in. All three gates now set their threshold by the same rule, so the board isolates judging quality and the clock. Online learning is separated from the judging. You're right that the deployable label is a delayed outcome (a line cleared a few moves later); we'll treat that on its own.

Laya's measured latency (0.015 s) is wired in too. On the updated board, over 1,000 paired matches in the same 60 s, ZTC clears 3.6 more lines than JEV and 0.9 more than Laya.

The page is live with the changes. Give it a run.

Ran the new board against its source. The shared rule holds, THR = logit(0.559)/SEP, and the delayed label (no new hole, a clear within 5 moves) answers my question.

One thing moved under the rename. At AUC 0.50, SEP = 0 and THR = Infinity. Laya no longer reads a score. It re-rolls every move.

That makes the Laya board the null arm, and a useful one. p* = 0.38/(0.38+0.30) = 0.559 is also where the re-draw chain settles, while the first proposal is good at 0.50. A re-roll alone buys lift.

Closed form, P(good move kept) per board:
no gate 0.500
always re-roll (Laya) 0.540
ZTC 0.579
JEV 0.580
perfect gate 0.690

So about half of ZTC's +1.82 over no gate is the re-roll, not the judging. The judging value is ZTC minus always-re-roll, which is your +0.89 [0.66, 1.12]. Worth labelling that board as the null, since "accuracy pending" reads like a scorer.

A question on the sim itself: why is a second proposal better than the first? In an agent loop, should a re-sample beat the first sample, or should the first draw also sit at 0.559?

·

Good catch on the first draw. It now sits at 0.559, the rate the re-draw chain settles at, so a re-roll without judging buys nothing. At its measured 0.50 the Laya board is the always-re-roll null arm, and the page labels it that way.

Re-run on the updated board, 1,000 paired matches, same 60 s: ZTC minus always-re-roll +1.06 lines [0.82, 1.30], ZTC minus JEV +4.16 [4.01, 4.33]. Same moves: ZTC and JEV tie.

On your question: a re-sample shouldn't beat the first sample on its own. The measured 38% / 30% transition rates imply the first draw also sits at 0.559, which is what the board uses now.