Skip to content

Methodology

Rules first. Results second.

A benchmark is only as useful as its protocol. We separate generation, automatic controls, and structured review, and make exceptions visible.

Internal, versioned benchmark · Updated July 28, 2026. These results describe this test run. They do not guarantee the same outcome for every text, language, or future product version.

The protocol in six steps

  1. 01

    Fixed source texts

    Every provider receives the same source. This run contains ten cases with ordinary and structured citations.

  2. 02

    Production mode

    This archive records the Certentis production pipeline used on the test date. It is not a measurement of every later product version.

  3. 03

    Blinded quality review

    Rewrites are reviewed without provider labels against a fixed eight-dimension rubric.

  4. 04

    Separate semantic review

    Facts, qualifications, causality, and intent are checked against the corresponding source text.

  5. 05

    Hard structural controls

    Language, word ratio, and complete citation occurrences are recorded as separate metrics.

  6. 06

    Failures remain visible

    Missing outputs are not silently replaced. RewriteAI therefore has nine cases instead of ten.

Eight quality-review dimensions

  1. 01Grammar and language control
  2. 02Clarity and readability
  3. 03Coherence and structure
  4. 04Academic style
  5. 05Precision
  6. 06Naturalness
  7. 07Citation integration
  8. 08Semantic fidelity to the source

How to read the metrics

01

Quality · 0 to 100

The structured total summarizes the review dimensions while semantic fidelity remains visible separately.

02

Citations · exact or changed

We record whether every complete citation occurrence is preserved byte for byte in the output.

03

Word ratio · output ÷ input

1.00 means equal word counts. The ratio measures length, not content quality.

04

External AI classification

Human and Mixed are added as an orientation value. It is a model probability, not a sentence share or proof.

Limitations and next steps

  1. 01Ten cases form a useful product test, but not a complete representation of every text type.
  2. 02The published material does not yet identify the reviewers or their configuration, define the complete scoring rubric and weighting, or provide case-level assessments. The aggregate scores cannot therefore be independently reconstructed from this page.
  3. 03The benchmark’s meaning review is separate from the live product: automatic structural checks do not establish semantic equivalence. Review the final text yourself.
  4. 04Source texts cannot be published in full where licensing or privacy restricts redistribution.
  5. 05External models and competitor products may change after the test date.
  6. 06We will publish accuracy figures for the AI Detector and Plagiarism Checker only after testing against suitable labeled reference corpora.

Certentis Evidence

Evidence before promises.

Results, test rules, and limitations are kept separate so it remains clear what was measured and what was not.