Quality · 0 to 100
The structured total summarizes the review dimensions while semantic fidelity remains visible separately.
Methodology
A benchmark is only as useful as its protocol. We separate generation, automatic controls, and structured review, and make exceptions visible.
Internal, versioned benchmark · Updated July 28, 2026. These results describe this test run. They do not guarantee the same outcome for every text, language, or future product version.


Every provider receives the same source. This run contains ten cases with ordinary and structured citations.
Certentis is tested with the pipeline actually used in production, not a selectively optimized demo configuration.
Rewrites are reviewed without provider labels against a fixed eight-dimension rubric.
Facts, qualifications, causality, and intent are checked against the corresponding source text.
Language, word ratio, and complete citation occurrences are recorded as separate metrics.
Missing outputs are not silently replaced. RewriteAI therefore has nine cases instead of ten.
The structured total summarizes the review dimensions while semantic fidelity remains visible separately.
We record whether every complete citation occurrence is preserved byte for byte in the output.
1.00 means equal word counts. The ratio measures length, not content quality.
Human and Mixed are added as an orientation value. It is a model probability, not a sentence share or proof.
Certentis Evidence
Results, test rules, and limitations are kept separate so it remains clear what was measured and what was not.
