AI humanizers compared directly

Certentis vs StealthGPT Heavy

How does StealthGPT Heavy perform in an academic AI humanizer comparison? The archived run also selected PhD, but StealthGPT's current API documentation says that tone is ignored when rewriting existing text.

Internal benchmark · July 28, 2026. The same source texts, blinded review, and separate checks for meaning, language, length, and citations.

Overall quality

85.4Certentis
70.0StealthGPT Heavy PhD
Certentis vs StealthGPT Heavy

Product profile and terms

What both products actually offer

The benchmark measures output quality. This section adds the current product context: modes, limits, languages, platforms, and publicly documented safeguards.

Certentis in everyday use

Certentis

A controlled humanizer for academic and professional texts, with separate AI analysis and a traceable history.

Certentis does not optimize for a scanner classification alone. The existing language, length, meaning, and citations are treated as separate quality conditions.

Product profile

StealthGPT Heavy PhD

A detector-oriented writing suite with a humanizer, writer, and research agent. Heavy is the humanizer model; PhD is a tone for newly generated essays and, according to the API, is ignored for rewrites.

StealthGPT focuses on substantial rewriting and high volumes. Its own test shows that formatted citations can work well, while structured citeproc cases require particularly careful review.

Features and terms compared

Price and trial access

Certentis

Word credits; only the input text counts. Internal retries within one process do not consume additional words.

StealthGPT Heavy PhD

Pricing page: Essential $1.00, Pro $1.45, Business $2.15, and Enterprise $7.15 per day as the displayed plan equivalents.

Input limit

Certentis

Up to 30,000 characters and 3,000 words per input in the current web app.

StealthGPT Heavy PhD

Depending on the plan, 1,000, 1,500, 2,000, or 20,000 words per request; the API still recommends short outputs for optimal results.

Languages

Certentis

Automatic language detection; the rewrite remains in the original's dominant language.

StealthGPT Heavy PhD

The current pricing page advertises 100+ languages; other official pages give different figures.

Modes and workflow

Certentis

One controlled mode with local checks followed by an analysis report.

StealthGPT Heavy PhD

Heavy or Lite plus Quality or Fast mode. PhD applies to newly generated essays, not to the humanization of existing text.

Academic texts and citations

Certentis

Existing citation blocks are protected and checked for completeness after the rewrite.

StealthGPT Heavy PhD

Writer and Agent provide separate research and citation features. Byte-exact protection in the Heavy humanizer is not documented.

Platforms and API

Certentis

Web app with history, AI analysis, and plagiarism checking in the same workspace.

StealthGPT Heavy PhD

Web suite and API. The API costs $0.20 per 1,000 calculated words and counts both input and output.

The site offers weekly, monthly, and annual plans, but the table displays daily equivalents. According to the pricing table, consumer subscriptions do not include API access.

01

Certentis · Strengths of the offering

  • Quality, meaning, language, length, and citations are checked separately.
  • The report separates text changes from AI signals and makes both traceable.
  • The archived comparison uses identical inputs and a blinded review.
02

StealthGPT Heavy PhD · Strengths of the offering

  • Explicit modes for academic and professional texts.
  • Strong results for formatted citations in the archived test.
  • A broad selection of writing and rewriting tools.
03

What to consider before buying

  • According to the API, PhD is not an active setting when rewriting existing text.
  • Official statements about the number of supported languages conflict.
  • No traceable citation lock for citeproc keys is described.

Official product sources

Provider information verified on August 3, 2026

Prices and features may change. Provider claims were not treated as evidence of quality; only the separate Certentis benchmark is used for that assessment.

The result in one sentence

+15.4 quality points ahead

StealthGPT maintained the language and many citation sets, but shortened texts substantially. Certentis led clearly on overall quality, semantic fidelity, and Human-plus-Mixed.

Quality scale 40–100 · truncated axis

Certentis85.4
StealthGPT Heavy PhD70.0
405060708090100
Certentis vs StealthGPT Heavy: Overall quality

Semantic fidelity

91.7Certentis
72.8StealthGPT Heavy PhD

Exact citations

8/10Certentis
6/10StealthGPT Heavy PhD

Distance from source quality

-8.9Certentis
-24.3StealthGPT Heavy PhD

Best version in blinded review

4/10Certentis
2/10StealthGPT Heavy PhD

The comparison at a glance

Every metric in the head-to-head test

A single AI detector score cannot decide the result. A useful rewrite must also remain natural, faithful to the source, stable in language, and dependable in its citations.

MetricCertentisStealthGPT Heavy PhDStronger result
Overall quality85.470.0Certentis
Semantic fidelity91.772.8Certentis
Exact citations8/106/10Certentis
Correct language10/1010/10Tie
Word ratio0.980.82Certentis
Human + Mixed92.96%68.75%Certentis

Certentis: 10 usable cases · StealthGPT Heavy PhD: 10 usable cases

Overall quality

Eight quality dimensions, not one headline score

The blinded review scored every output separately. These profiles show whether a tool merely sounds fluent or also preserves structure, precision, evidence, and meaning.

StealthGPT produced an unusually split profile: very strong grammar and academic style, but markedly weaker coherence and semantic fidelity across the full run.

Grammar & language93.9
StealthGPT Heavy PhD
91.6
Clarity & readability91.8
StealthGPT Heavy PhD
82.6
Coherence & structure89.4
StealthGPT Heavy PhD
73.0
Academic style80.3
StealthGPT Heavy PhD
85.9
Precision84.6
StealthGPT Heavy PhD
78.0
Naturalness87.5
StealthGPT Heavy PhD
79.8
Citation integration94.2
StealthGPT Heavy PhD
82.0
Semantic fidelity91.7
StealthGPT Heavy PhD
72.8
Eight quality dimensions, not one headline score: Certentis, StealthGPT Heavy PhD

Citeproc keys / Formatted citations

Citeproc and formatted citations measured separately

Five source texts used structured citeproc keys and five used already formatted references. That prevents format-specific weaknesses from disappearing inside an average.

On formatted citations, StealthGPT reached 86.2 quality points and narrowly exceeded Certentis. In citeproc, incomplete and heavily shortened outputs pulled quality down to 53.8.

Citeproc and formatted citations measured separately

Citeproc keys

MetricCertentisStealthGPT Heavy PhD
Quality85.053.8
Fidelity90.656.0
Citation integration92.870.2
Citations exact3/52/5
Human + Mixed97.4%69.53%
Word ratio0.960.63

Formatted citations

MetricCertentisStealthGPT Heavy PhD
Quality85.886.2
Fidelity92.889.6
Citation integration95.693.8
Citations exact5/54/5
Human + Mixed88.51%67.98%
Word ratio0.991.02

Observation from the benchmark

What the blinded review actually observed

These observations come from individual cases in the archived run. They explain the metrics, but they are not blanket claims about every future output.

01

Very strong formatted rewrites

Several outputs with standard references were fluent, terminologically precise, and dependable in their citation handling.

02

Some citeproc outputs stopped early

Large parts were missing in the German and Spanish citeproc cases; average length for this suite was only 63 percent of the source.

03

Isolated foreign characters and term errors

The review found foreign script fragments or inaccurate terminology in individual cases, even though the target language check passed.

What the numbers mean in practice

Which AI humanizer fits which priority?

The right tool depends on whether you prioritize the highest detector classification, a particular writing style, or a controlled academic rewrite.

Certentis

Certentis is the better fit when

different citation types must be processed in one stable workflow without major shortening.

StealthGPT Heavy PhD

StealthGPT may fit when

the document mainly contains formatted citations and every output will be checked manually for completeness.

01

The 0.82 average word ratio means StealthGPT shortened the texts by about 18 percent on average.

02

Shorter outputs are not inherently worse, but in this run they coincided with lower semantic fidelity.

03

For academic text, preserving intended meaning and formal evidence matters more than maximizing the amount of rewriting.

Internal benchmark

How the tools were compared

Every provider received the same source texts. Outputs were scored against a fixed rubric without visible provider labels. Missing results, language drift, and changed citations remained visible.

How the tools were compared
  1. 01

    Ten identical source texts with standard and structured citations

  2. 02

    Blinded quality review across eight dimensions

  3. 03

    Separate comparison of facts, qualifications, and intended meaning

  4. 04

    Additional checks for language, word ratio, and complete citations

Review the methodology

Limits of this comparison

  • 01The results describe a dated run and do not guarantee every future output.
  • 02Provider models, prices, and product versions may change after the test.
  • 03Ten cases are a practical product test, not a complete representation of every text type.
  • 04Human + Mixed is an external document classification, not a percentage of human-written sentences.

Frequently asked questions

Is Certentis a StealthGPT alternative?

Yes. Both produce more natural rewrites. Certentis additionally focuses on controlled preservation of meaning, language, and citations.

Why was StealthGPT shorter in the test?

The measured outputs averaged 82 percent of the input word count. The benchmark describes that effect without claiming to know the competitor's internal cause.

Which provider preserved more citations exactly?

Certentis preserved eight of ten complete citation sets; StealthGPT preserved six of ten.

Can one AI humanizer test settle which provider is best?

No. This run is a dated sample of ten demanding multilingual texts. It reveals concrete strengths and risks, but it does not replace testing your own document.

Why is Human + Mixed not used as the only winning metric?

The classification describes how an external scanner categorizes a document. It does not assess factual correctness, citation preservation, language stability, or argument quality.

What counts as an exact citation set?

Every reference present in the source had to remain complete and unchanged in the rewrite. A missing, reformatted, or damaged citation meant the case did not pass as exact.