$ ls evidence/
evidence/
does it work? answered from our own runs, graded on the strength of the evidence
limited evidenceDoes an AI code reviewer catch real problems?The review stages logged 135 coded findings across 8 runs, 31 of them BLOCKER and 92 coded FIXED, plus 25 cold-review escalations and 5 proofreader holds; nothing here shows what share of real problems that was.8 minmixed evidenceHow often does a critic loop change the output?A critic loop fired at least once in 49 of the 67 runs with a record, a 73% share, mostly at the architect and prompts stages, while only 23 of 61 runs recorded any rework spend at all and the median was 0.0.8 min