Every review package opens with a number out of 100. It is the most misread thing in the product, so this page says plainly what it measures — and what it does not.
The score rates your document, not the review. A low number does not mean the board failed, or that the reviewers were unsure. It means they read the artifact carefully and found serious problems with it. A document scoring 18 out of 100 has been taken seriously.
The number is grouped into five bands, and the band is shown next to the score in the product so the direction of the scale is never ambiguous.
The board found little needed change in this document.
Sound overall, with specific improvements worth making.
Real problems the board expects you to address before this goes further.
The board found serious problems with this document.
Severe problems. The document should not proceed as written.
Documents get submitted because a decision is pending, and drafts at that stage tend to carry unresolved risk — that is why you wanted a review. Reviewers are instructed never to soften a finding, so a well-written draft with real gaps scores low. Writing quality is not what is being measured.
The score earns its keep as a before-and-after measure. Revise against the recommendations, run the document again, and the movement between the two runs tells you more than either number alone.
Next to the score you will see TYPE MATCH — high, medium or low. That is how confidently your document was classified into a canonical artifact type: architecture decision record, threat model, privacy assessment, and so on. It says nothing about quality. It matters because the detected type decides which specialists the board seats.
Most documents report medium. The classifier requires an unusually strong signal before it will say high, which is deliberate — overstating certainty about a document type would quietly seat the wrong reviewers.
Usually not. Running the same document through a five-reviewer and a seven-reviewer board produced an identical score in our own testing. The larger board earned its price differently: it raised four issues the smaller board never mentioned, because it seats specialists the smaller panel does not include. More reviewers buy breadth of scrutiny, not a different verdict — which is worth knowing before you choose a panel size.
A panel of independent AI reviewers that read the same document in parallel, each from a different expert perspective, and return one combined review. Suggestibility seats three, five or seven reviewers depending on the plan, and preserves disagreement between them instead of averaging it into a single answer.
Asking one model the same question several times produces variations of one opinion, which reads like consensus without being it. Reviewers drawn from independent model families disagree for real reasons — different training, different blind spots — so their agreement carries information and their dissent is worth recording.
Yes. Every finding names the model and provider that produced it, and the synthesizer that merged the board is disclosed separately from the reviewers themselves, because it merges the board rather than sitting on it.
It is preserved as a minority position, attributed to the reviewer that held it, in its own section of the review package. Averaging dissent away would remove the main reason to convene several reviewers at all. See how this compares to a human board.
Compare a document to its own earlier revisions rather than to someone else's document. Each artifact is scored against the standards of its own type, so a threat model and a pricing policy both scoring 60 are not making the same statement.
Yes. Every completed review has a readable view alongside the full technical detail — plain English on three- and five-reviewer boards, and an executive summary on a seven-reviewer board. Both keep the severities and the recorded dissent; they change the register, not the findings.