- Pangram produces contradictory results: the same text can be flagged as 100% AI or 100% human depending on context, creating a 'Russian doll' of inconsistent outcomes.
- The percentage meter is misleading because many users interpret '100% AI' as high confidence, but it actually indicates the portion of text suspected to be AI-generated; the meter also frequently outputs 100% for hybrid texts.
- False positives are easy to induce: the author wrote his own text that Pangram flagged as 100% AI with high confidence, and AI-generated text can be made to appear human by embedding it in human writing.
- Even a low error rate (e.g., 1 in 10,000) leads to many false accusations when applied to large volumes of text, especially every paragraph becomes a separate test.
- The tool's design encourages a culture of 'gotcha' accusations among writers, exploiting professional jealousy and resentment rather than providing reliable evidence.
- Pangram's founder acknowledged the issue with short text segments and the inevitability of adversarial false positives, but the author argues for epistemological modesty and warns against treating any detection tool as a final authority.