Pangram is an AI text detector that many consider one of the strongest options available today. But confidence in this assessment requires some evidence because no amount of vendor assurances can be considered complete proof.
This review keeps three things separate:
- What Pangram claims about itself
- What independent research has actually found
- What happened when we ran real samples through it ourselves
Using a single copy-pasted ChatGPT paragraph wouldn’t have provided enough information by itself. That’s why we tested the Pangram AI detector across an academic essay, creative writing, and a social media post to evaluate its detection capabilities.
Here’s what we found.
What Is Pangram AI Detector?

Pangram is an AI text detector developed by Pangram Labs to determine how much of a piece of writing was written by a person or an AI model and classifies it into one of three buckets:
- Fully human-written
- AI-assisted or mixed authorship
- Fully AI-generated
It’s used across a few different high-stakes settings:
- Education, for checking student submissions
- Publishing and media, for screening manuscripts and articles
- Content moderation, for flagging AI-generated posts at scale
- Recruiting, for reviewing written work samples
- Professional organizations, for verifying originality in submitted content
Pangram 4 vs. Older Pangram Versions
The newest version of Pangram’s production model was released as Pangram 4 on July 29, 2026. It fixes real problems: how it treats humanized text and mixed-authorship writing.
If you have read older reviews of Pangram, they probably were written about some version other than Pangram 4. It will continue to be available until September 30, 2026, but this review is built entirely around Pangram 4, the version most people are using today.
How Does the Pangram AI Detector Work?

The Pangram AI detector is a classifier. It has been trained on large collections of both verified human writing and verified AI writing, and it uses these to learn larger, more complex patterns that separate the two.
In accordance with the methodology published by Pangram, training involves:
- Hard-negative mining, where the model is repeatedly shown its own mistakes to sharpen its edges
- Synthetic, mirrored AI examples built to match human writing topic for topic.
- Iterative retraining as new AI models and writing styles emerge.
Pangram 4 also adds sentence-level output, so instead of one blanket score, you can see which parts of a document read as human, AI-assisted, or fully AI-generated. Pangram describes the model as a mixture-of-experts backbone with three separate task heads, one for provenance, one for mixed authorship, and one for humanizer detection.
On its own benchmarks, Pangram reports a 0.0041% false-positive rate and a 0.3396% false-negative rate for Pangram 4. Those are Pangram’s own internal results, generated without outside verification, but they give a sense of how tightly the model is tuned to avoid both flagging clean human writing and missing real AI content.
Does Pangram Just Look for Predictable Words?
No. Traditional techniques were based on perplexity, which was simply a way to identify whether words in the writing had too many “predictable” types of word choice. These methods fail to identify many cases with no legitimate basis and therefore produce a large number of false positive identifications.
Pangram evaluates patterns across an entire sample, looking well beyond any single em dash, transitional phrase, or sentence structure.
Can Pangram Detect Humanized or Paraphrased AI?
Humanizers strip out the artificial-sounding language detectors rely on, which makes AI text harder to catch.
Pangram 4 has a dedicated humanizer-detection neural network designed specifically for this type of challenge. According to Pangram 4, they report a detection rate of approximately 98.83% when detecting the presence of AI-generated content using an internal benchmark covering 13 different types of commercially available AI-humanizing software.
That detection rate is worth paying attention to, and it comes from Pangram’s own internal testing benchmark. If you’re curious how a detect-and-fix workflow handles this same problem, Walter’s AI humanizer is worth a look.
Pangram Features, Integrations, and Pricing in 2026
| Feature | Available? | Why it matters |
|---|---|---|
| Sentence-level scoring | Yes | See exactly which parts of a document are flagged |
| Humanizer/mixed-authorship detection | Yes | Catches paraphrased or “humanized” AI text |
| Image detection | Yes | Separate feature, weaker than its text detection |
| File upload and OCR | Yes | Scan documents and scanned/image-based text |
| Google Docs integration | Yes | Scan directly inside your writing workflow |
| Chrome extension | Yes | Scan content while browsing |
| 20+ languages | Yes | Multilingual detection, not English-only |
| API and institutional/LMS options | Yes | Built for schools, publishers, and platforms |
Is Pangram Free?

Yes, but with some restrictions. On the free plan, there is a daily limit of 20 credits. Each credit allows for 100 words of scanned text. So on a daily basis, regardless of how many times you run a scan using the free service, you will be limited to 2,000 words of total scanned text.
Paid tiers raise that ceiling: Individual runs for 300,000 words at $15/month if billed annually. Professional runs for $45/month for 1.5 million words if billed annually.
While their website says “3 scans per day”, these are related to image detection and not text scanning. Text scans are governed by the 2,000 word/day credit cap instead.
What Content Works Best?
According to Pangram’s own documentation, Pangram 4 is best suited for natural-language samples with a minimum of 50 words. The following types of content are less suitable for Pangram:
- Short one-line replies
- Source code
- Reference lists and citations
- Heavily templated text
- Technical manuals and mathematical content
If you have a full-length essay, article, or manuscript that needs to be scanned, then this is where you will find yourself using Pangram in its primary use case.
Pangram AI Detector Stress Test (2026): Real-World Scenarios Pangram Must Survive
Most detector reviews paste in one raw ChatGPT paragraph into the tool, run it through once, and provide you with just one number. That tells you very little about what happens when an author uses a tool in real-world writing situations they encounter.
So we used a formal testing structure instead. It included verified AI-generated content and verified human-written content across three real-world content types.
Pangram AI Detector Results Table
| Use Case | AI-Generated | Human-Written |
|---|---|---|
| Academic Essay | 100% AI ✅ | 0% AI ✅ |
| Creative Writing | 100% AI ✅ | 6% AI ✅ |
| Social Media Post | 100% AI ✅ | 0% AI ✅ |
| Average | 100% AI | 2% AI |
Test Setup: The Three Real-World Use Cases
We tested Pangram across the standard three use cases that we use for evaluating every detector so results stay comparable across tools:
- Academic essay, the highest-risk case for both students and educators
- Creative writing, a style-based content type that can be difficult to identify when written by AI
- Social media post, a very short piece of content in which detection can be unpredictable
All samples were left unedited before running each test.
Test Samples Used
Below are all 6 samples used during the testing.
AI Academic Essay vs Human Academic Essay

AI Creative Writing vs Human Creative Writing

AI LinkedIn Post vs Human Social Media Post

Link to these samples:
Academic Essay: AI (Artificial Intelligence and the Redesign of Assessment in Higher Education) and Human (Is Google Making Us Stupid?)
Creative Writing: AI (Student’s Nightmare) and Human (The Gift of the Magi)
Social Media Post: AI (Transforming Content Marketing) and Human (Is It OK to Publish Late)
Pangram Results: AI and Human Academic Essay


Pangram Results: AI and Human Creative Writing


Pangram Results: AI and Human Social Media Post


What the Pangram Results Show

Raw AI detection: Perfect across all three raw AI-generated samples (academic, creative, social). No missed signals.
Human writing / false positives: Low rates across the board. Only the 6% in human-generated creative writing had a positive result. And even this is well under any reasonable flagging threshold.
Short-form content: This type of content can break a detector, but the 100% results from the social post were maintained by the Pangram AI detector with no drop-off.
Consistency: In all three raw AI samples tested, the Pangram detector remained consistent and reported 100%.
One caution worth stating plainly: six samples is a controlled test. These results show how well the Pangram detector worked on those specific samples. We do not believe these results will reflect every sample of writing you may send to the detector.
Pangram vs. Walter Writes: Same Samples, Same Conditions
To see how these two tools perform on the same test data samples, we ran all 6 of our original samples through Walter Writes without editing anything.
Side-by-Side Results Table
| Sample Type | Use Case | Pangram Score | Walter Score | Better Result |
|---|---|---|---|---|
| AI-Generated | Academic Essay | 100% AI ✅ | 99% AI ✅ | Both |
| Human | Academic Essay | 0% AI ✅ | 5% AI ✅ | Both |
| AI-Generated | Creative Writing | 100% AI ✅ | 99% AI ✅ | Both |
| Human | Creative Writing | 6% AI ✅ | 1% AI ✅ | Both |
| AI-Generated | Social Media Post | 100% AI ✅ | 99% AI ✅ | Both |
| Human | Social Media Post | 0% AI ✅ | 5% AI ✅ | Both |
Walter Writes Results: AI and Human Academic Essays


Walter Writes Results: AI and Human Creative Writing


Walter Writes Results: AI and Human Social Media Post


Detection Accuracy
In terms of detection accuracy, both tools have performed very well. Pangram was able to detect each of the AI-generated samples as 100% AI, and Walter landed at 99% on all three. That’s about as close as two tools can score, and either result will provide educators or publishers a reliable and meaningful way to use AI detection tools in their work.
False Positives and False Negatives
No false positives occurred with either tool. All human-written samples had similar scores with both tools, ranging from 0% to 6%. No human-created content was falsely identified as AI-created.
Performance on Short-Form Content
Short-form content is often where detectors lose accuracy, but both tools held steady here. On the social post, Pangram achieved a score of 100%, while Walter Writes achieved a score of 99%. No significant performance loss was observed by either tool when comparing their scores to those they obtained using the longer samples.
Workflow Difference: Detector vs. All-in-One Writing Platform
Based on these results, Pangram and Walter sit in the same performance tier. Both are strong, reliable choices for catching raw AI content.
Where they differ is what happens after you get a score. Pangram is a detection product, so it tells you what a piece of text likely is and stops there. Walter combines detection with humanization and writing tools in the same editor, so a flagged result can be revised in the same place it was found.
Try Walter’s AI detector on your own writing to see how it scores.
Is Pangram AI Detector Accurate?
In short: yes, according to our own stress-testing, but “accurate” isn’t a single number.
Pangram reports very strong internal benchmarks. Independent testing has also produced strong results. A 2025 Becker Friedman Institute working paper from the University of Chicago tested Pangram alongside GPTZero and Originality AI. In the researchers’ test set, Pangram produced essentially no false positives on medium- and long-form passages and kept its false-positive rate below 1% on short passages. Its false-negative rate was also near zero on medium- and long-form AI text.
This is more credible than Pangrams’ internal testing because external testing by independent researchers was done here. But they say that the results are valid only for the dataset upon which this study was done. There can be significant variations across real-world writing, so these results may vary significantly in other real-world contexts.
Accuracy can still be based on many factors. There are many ways that the results can vary, such as false positives, false negatives, the amount of content being evaluated, and how the material was created.
Pangram has also been at the center of real controversy. In one widely reported case, Pangram’s own report on the novel Shy Girl came back at 78.3% AI-generated, a result that led to major fallout for the author.

We break down that full accuracy debate, including where detectors get it wrong and how accurate the Pangram AI detector really is.
Is Pangram on Substack?
Yes. The integration of Substack with Pangram-powered AI detection took place in 2026. In this way, writers and readers can identify the use of AI directly within the platform.
Subsequently, there was quite a bit of debate surrounding this addition. Especially the writers who use AI writing tools as an integral part of their editing process were not happy about Substack implementing this technology.
We cover the integration, the backlash, and what it means for writers publishing on Substack in Substack’s Pangram AI detector.
Pangram Pros and Cons

Pros
- Strong track record in independent research, backed by researchers outside the company
- Detailed, sentence-level output showing exactly which parts of a document read as human, AI-assisted, or AI-generated
- Perfect raw AI detection in our own stress test, with no false positives on human-written content.
- Broad integration options available to users, including Google Docs, Chrome browser extension, API access, and institutional/LMS option.
- Multilingual support exists across 20+ languages.
- A usable free tier, at 2,000 words of text scanning per day
Cons
- A score estimates likelihood. Determining the quality of an article cannot be measured through a percentage alone, even though we tested Pangram and concluded it was highly accurate in terms of its ability to measure AI-written material.
- Not built for every content type. It does not perform well when used on very short responses like tweets, blocks of code, citations/references, or on heavily templated writing.
- High-stakes use needs context. Users should always consider the risks involved with creating a single score for high-stakes uses. In most educational institutions, one number may result in serious consequences.
- Detection-only workflow. Once it determines the likelihood of what a text probably is, Pangram stops there. It does not assist in editing the flagged items in the same location.
Who Should Use Pangram and Who Is Better Served by Walter Writes?

Pangram Makes Sense For
- Universities and schools already building AI detection into their academic integrity workflows.
- Publishers and content teams that need API or LMS-style screening at scale
- Anyone whose main need is classification and investigation, not revision
Walter Writes Makes More Sense For
- Writers, marketers, and creators who need to detect and improve text in the same place
- Anyone who doesn’t want a standalone detector followed by a separate rewriting or humanizing tool
- Readers who want to fix a flagged result in the same place they found it
If you are looking for just a score, Pangram will give you a serious number. If you want to act upon that number also, then using a process of identifying and humanizing your findings will allow you to accomplish this within one application.
Pangram AI Detector Review: Final Verdict
Pangram earns its reputation. The real independent research supports it, and our testing also showed that it detected raw AI-generated content and never flagged legitimate human content.
A percentage is merely an approximation. Whether the tool reports 78% or 100%, that distinction matters most when careers, reputations, or academic standing are at risk.
If you are interested in finding a classifier only, then Pangram can be a good option, as it appears to be a reliable option available to accomplish this task.
But for readers who need to identify and correct problems without changing tools, AI detection with Walter Writes provides a better overall experience due to its combination of detecting AI and humanizing it, plus writing capabilities within the same environment.

Frequently Asked Questions About the Pangram AI Detector
Is Pangram a Good AI Detector?
Yes. Independent research has shown that Pangram has been successful at detecting AI-generated content using other detectors, and when we ran our stress tests, we were able to detect all samples of raw AI content without generating a single false positive on human-written content.
How Do You Beat Pangram AI Detection?
While no detection method is foolproof, Pangram 4 was designed to detect both humanized and AI-generated content, as well as content that had been rewritten or paraphrased by humans. See how AI humanization tools actually perform against detection.
Is Pangram Better Than GPTZero?
Pangram performed better than GPTZero on raw AI detection for our identical data sample, especially on the short form of content. You can refer to this GPTZero review and stress test for a better understanding.
Is the Pangram AI Detector More Accurate Than Turnitin?
We only compare tools where we have direct, identical testing. We cannot say if the Pangram AI detector is more accurate than Turnitin based on our current data because we do not have Turnitin tested using these same test methods.
How Accurate Is the Pangram AI Detector?
The accuracy of Pangram AI will depend upon what type of content is being detected, how long the content is, and if the content has been paraphrased or otherwise humanized. The full breakdown of the accuracy of the Pangram AI detector can be found in this guide.
Is the Pangram AI Detector Free to Use?
Yes, with a daily limit of 2,000 words of text scanning, based on 20 credits per day at 100 words per credit.
Can Pangram Be Bypassed or Tricked?
There isn’t a detector that guarantees to detect every single deception. Pangram 4 was developed with both humanized and adversarial text in mind, but results are probability-based rather than guaranteed.
Does Pangram Detect ChatGPT, Claude, and Gemini Equally Well?
Pangram is a tool that uses benchmark results for comparing major AI model families, though performance will vary depending on the individual model and content type and does not have to be perfectly identical when comparing all models.
Is Pangram used by universities and publishers?
Yes. Pangram has created institutional and LMS integration options, and it’s been used in several university and academic settings as well as within the publishing industry, including some high-profile instances covered in this article: Is Pangram Accurate?
Other AI Detector Reviews
- Scribbr AI Detector review
- Phrasly review
- Grubby AI review
- JustDone AI review
- TwainGPT review
- What ChatGPT Zero is
EEAT Score and Link:
https://chatgpt.com/share/e/6a9e721b-18d8-83e8-803a-571b814aea0b



