What the allegation rests on

GPTZero examined four PwC Middle East reports published between 2024 and 2026 and reported a pattern it calls Vibe Citing: references cited loosely, often without correct titles, URLs or authors, and not supporting the claims they are attached to. The Financial Times reported it first.[1]

The same review scored one report, Transforming Governance, as 84 percent likely to be entirely AI-generated. That proportion estimates statistical resemblance; it does not name who or what produced the text. The verifiable half of the allegation is the other one: a reference either supports the sentence it is attached to or it does not, and anyone who opens the reports can see which.[1]

The detector's own margin of error

This week Pangram raised 9 million dollars and reports over 99 percent accuracy at detecting AI-assisted and mixed writing, with roughly one in 10,000 human-written documents wrongly flagged. Both figures are the company's own and both carry the same constraint: at the scale of a consultancy's back catalogue even that rate produces false accusations, and the report describes no appeal route for the person wrongly flagged.[2]

On PwC's side the strongest competing account is this: careless citation under deadline leaves the same trace as a fabricating model. The firm said it takes accuracy seriously and is “updating a limited number of supporting citations”, without explaining how the errors occurred. That reply concedes the citation problem and leaves the question of how the text was written where it was.[1]

What can be checked today and what cannot

GPTZero says it previously found the same problems in KPMG, Deloitte and Ernst & Young reports, which widens the pattern claim without changing the kind of evidence behind it. Listing the corrected references per report, original and replacement side by side, and stating whether the underlying claim survives, is something that could be done today for all four reports. The 84 percent is not resolved by that list: the authorship question stays at the level of an allegation while the detector's threshold and error rates go undisclosed.[1], [2]