Skip to content
Academic Integrity22 min read

Why Single Aggregate AI Scores Fail: Sentence-Level Stylometry and Defensible Pre-Submission Audits

Two manuscripts can share the same 30% AI score and need opposite responses. This guide explains why whole-document percentages conceal the evidence editors actually assess, and sets out a six-step audit that works at sentence level.

Why Single Aggregate AI Scores Fail: Sentence-Level Stylometry and Defensible Pre-Submission Audits

A single aggregate AI score is an inadequate basis for judging the integrity of a scholarly manuscript. It reduces a complex document — containing human drafting, discipline-standard phrasing, citations, quotations, editing, collaboration, and possibly AI-assisted text — to one opaque percentage. It cannot identify the sentence that requires review, explain the basis of the flag, distinguish raw AI text from AI-paraphrased writing, or tell an author what corrective action is justified.

For PhD candidates, postdoctoral researchers, faculty members, and laboratory leads, the practical goal is not to "beat" an AI detector. It is to produce a submission record that is accurate, attributable, transparent, and defensible if scrutinized by an editor, reviewer, university ethics committee, or research-integrity office. That requires a different pre-submission model: inspect sentences in context, verify source relationships at DOI level, preserve drafting provenance, and document human editorial judgment.

Why a single score cannot answer an editor's question

The real problem is not detection — it is evidence

AI tools have changed manuscript preparation without changing the central conditions of authorship. Human researchers remain responsible for the research question, methods, data provenance, analysis, interpretation, citations, conclusions, and final approval. That responsibility persists whether AI is used for grammar correction, translation, literature organization, code support, restructuring, or content generation.

The challenge arises when researchers confuse a numerical detector output with a meaningful integrity assessment. A manuscript-level score may be convenient, but scholarly publishing is not assessed at that level of abstraction. Editors evaluate concrete passages. They ask whether a claim is supported, whether a source is accurately cited, whether the language misrepresents the data, whether authorship is authentic, and whether the authors have complied with a journal's AI and publication-ethics policies.

An aggregate score cannot resolve these questions because it conceals heterogeneity. It may combine a small number of highly suspicious sentences with many unremarkable sentences, producing a moderate percentage. Alternatively, it may assign a high score to a paper whose methods language is conventional, whose English has been professionally edited, or whose author writes in a highly regular style. The percentage alone does not reveal which interpretation is correct.

Two manuscripts, one score, opposite problems

Consider two manuscripts, each assigned the same 30% aggregate AI-likelihood score.

  • In the first manuscript, two paragraphs of the Discussion were inserted from unreviewed chatbot output. They contain generic claims, unsupported causal language, and references that have not been verified.

  • In the second manuscript, a non-native English-speaking author has used grammar assistance on an otherwise human-authored paper. Several routine sentences in the abstract and methods use predictable syntax, while all claims are supported by data and citations.

The numerical result is identical; the editorial response should not be. The first manuscript requires substantive verification, correction, possible disclosure, and a review of the team's AI-use process. The second may require no textual change at all, or only a proportionate disclosure depending on the target journal's policy. A whole-document score cannot make that distinction.

This is a structural limitation, not simply a limitation of a particular platform. AI-detection methods infer likelihood from linguistic and statistical features. Stylometry may analyse function-word distributions, lexical diversity, sentence length, punctuation patterns, readability, syntactic regularity, and predictability. Such features may be informative, but none constitutes direct proof of authorship or misconduct. They can be influenced by disciplinary convention, text length, translation, copy-editing, collaborative writing, accessibility tools, and the rapid evolution of AI models.

A detector result should therefore be treated as a review signal. It should initiate examination, not end it.

Why blind paraphrasing compounds risk

When authors see a high aggregate score, an understandable impulse is to "humanize" the text: vary sentence length, insert informal phrasing, reorder clauses, or run passages through another paraphrasing service. This response is methodologically weak and ethically hazardous. It changes surface form without addressing whether the sentence is true, appropriately attributed, or genuinely owned by the author.

Blind paraphrasing is particularly dangerous when it affects source material. Rewording a published finding, conceptual sequence, research rationale, or distinctive expression does not remove the need to cite the original work. If an author uses AI to disguise source dependence, a low textual-overlap result may conceal rather than solve an attribution problem. Similarity scores are evidence of matching text; they are not definitions of plagiarism. A manuscript can exhibit little verbatim overlap and still misappropriate another researcher's ideas through uncited close paraphrase.

The right question is not "How can we lower this score?" It is "What is the provenance of this sentence, what does it claim, and what evidence supports it?"

Why this matters at submission

Selective journals expect more than fluent prose. Their editorial processes evaluate scope, novelty, reporting quality, ethics declarations, authorship, conflicts of interest, image integrity, data transparency, citation practices, and textual overlap. COPE states that AI tools cannot be authors because they cannot accept responsibility for the submitted work; human authors remain accountable for material produced with such tools. Elsevier similarly emphasizes human oversight, verification, transparency, and protection of confidential manuscript inputs.

A manuscript that is scientifically sound may still receive intensified editorial scrutiny if it includes unexplained AI-generated passages, unreliable citations, unattributed source dependence, or weak documentation. A defensible audit does not guarantee acceptance, but it enables the team to show that it made reasoned, evidence-based decisions before submission. For the broader policy picture, see our guide to ethics and disclosure guidelines for AI in academic writing.

From aggregate detection to sentence-level review

A defensible workflow begins by changing the unit of analysis. Instead of asking whether an entire manuscript is "AI-written," ask whether each flagged sentence warrants further review. Sentence-level stylometry provides this granularity. It identifies passages whose linguistic characteristics merit inspection, then permits an author to compare those passages with the manuscript's logic, data, sources, drafts, and disclosed tools.

This distinction is decisive. A whole-document percentage is difficult to act on: it does not identify a location, an evidence type, or an appropriate revision. A sentence-level signal creates an audit target. Researchers can determine whether the sentence is a conventional disciplinary formulation, a generic machine-like assertion, a potentially AI-paraphrased rendering of a source, or a legitimate human-authored sentence whose provenance can be documented.

The SciCampus AI Detector structures this review with sentence-level probability colour bands:

  • Red (≥99%) — a highest-priority review signal. Check whether the sentence is raw AI-generated text, a repeated template, a close paraphrase, or a formulaic but defensible human sentence.

  • Orange (≥96%) — a strong signal requiring contextual verification, especially in abstracts, introductions, discussion sections, and transitions where generic AI phrasing can obscure the study's contribution.

  • Yellow (≥93%) — a lower-priority but still meaningful review prompt. Examine the sentence alongside its citations, drafting history, and surrounding argument before deciding whether any change is necessary.

The bands are not misconduct categories. Red does not mean "proven AI authorship," and yellow does not mean "safe." Their value is triage. They tell a researcher where to apply careful human editorial judgment first.

Raw output versus AI paraphrase

The central advantage of sentence-level inspection is not just localization; it is sub-classification. The detector can distinguish between probable raw AI-generated text and probable AI-paraphrased writing at sentence level. These classifications imply different risks and different responses.

Raw AI-generated text often appears as generic, fluent, high-certainty prose that is weakly connected to the study's actual design or findings. It may use broad claims — "These findings underscore the urgent need for a comprehensive response" — without identifying a sample, effect estimate, evidence boundary, or source. It may also contain invented references, overstated causal language, or fabricated methodological detail. The response is to verify every factual assertion against the study record, replace generalization with research-specific reporting, and disclose substantive use if required.

AI-paraphrased writing presents another problem. It may alter words and sentence order while preserving a source's conceptual structure, distinctive sequence of claims, or interpretive logic. The key issue is not whether the passage resembles chatbot prose; it is whether another scholar's contribution has been appropriated without appropriate attribution. The correct action is to locate the source, read the surrounding context, cite it accurately, and genuinely reframe the point in the author's own analytical voice. If the expression remains too close, quote it or remove it.

This sub-classification prevents two common errors. First, it stops authors from treating all flags as the same type of writing problem. Second, it stops them from treating a citation as a universal cure for textual similarity. Citation acknowledges an intellectual source; it does not automatically justify copying or near-copying expression.

DOI-grounded similarity is a separate evidentiary layer

Stylometry and similarity checking answer different questions. Stylometry asks whether linguistic features merit an authorship or provenance review. Similarity checking asks whether the manuscript overlaps with identifiable source material and whether that relationship is properly represented. Neither should be substituted for the other.

A generic similarity percentage has the same limitation as a generic AI score: it compresses the evidence. A 14% match may consist of references, a quotation, established methods language, thesis material, or multiple passages that need revision. Its meaning depends on source identity, location, quantity, wording, permissions, and attribution.

A search-grounded similarity index addresses this by connecting matches to DOI-linked sources where available. This enables the author to move from a matched sentence to the scholarly record, assess the actual source rather than a truncated snippet, and determine the correct disposition. For every substantive match, the team should classify it as one of the following:

  • Standard or unavoidable disciplinary language that does not create a meaningful attribution concern.

  • A correctly quoted and cited passage.

  • A properly attributed paraphrase that remains sufficiently independent in wording and structure.

  • Legitimate self-overlap requiring citation, disclosure, permission, or journal-specific revision.

  • A close paraphrase, unattributed idea, or copied expression that must be revised or removed.

A DOI-linked investigation produces a traceable editorial decision. It lets the author show not merely that a score was checked, but that the relevant source was located, reviewed, and resolved.

What modern publisher expectations require

Policies differ by publisher and journal, but their direction converges. AI cannot be listed as an author because authorship entails responsibility, conflict-of-interest disclosure, copyright obligations, and final approval — functions a tool cannot perform. COPE makes that accountability rationale explicit.

Elsevier states that authors may use AI tools to support manuscript preparation but must not use them as a substitute for human critical thinking, expertise, or evaluation. It expects authors to review AI output for accuracy, completeness, impartiality, fabricated references, privacy, and intellectual-property risks, and requires a declaration for substantive AI use; basic grammar, spelling, and punctuation checks are generally excepted.

IEEE requires authors to disclose AI-generated content, identify the system used, describe affected sections, and explain the level of use. It similarly distinguishes content generation from ordinary editing and grammar enhancement. The practical conclusion is clear: a responsible audit must identify not simply whether AI was used, but what it did, which passages or outputs it affected, and how human authors verified the result.

Audit the sentences that matter. SciCampus is free to start — no card required. Investigate sentence-level stylometry, distinguish raw AI output from AI-paraphrased writing, and review DOI-grounded similarity evidence before submission. Create a free account.

A six-step pre-submission audit

Step 1: Preserve a baseline before editorial intervention

Before conducting an AI or similarity review, preserve a dated, version-controlled copy of the candidate manuscript. Assemble supporting records: draft history, tracked changes, literature notes, reference library, protocol or preregistration, data-analysis outputs, figure source files, and author-contribution documentation. If AI was used substantively, create or complete a use log that records the tool, version, date, purpose, input class, affected section, human reviewer, and final disposition.

This baseline is vital. Without it, a team cannot readily show how a disputed passage evolved or distinguish original drafting from later editing. Do not manufacture a retrospective log. If prior use is uncertain, record that uncertainty honestly, consult co-authors, and ensure that the final disclosure is accurate.

Step 2: Run a sentence-level colour-band audit

Review the manuscript by probability band, beginning with red sentences, then orange, then yellow. Do not isolate sentences from their argument. Examine the preceding and following text, citations, figure and table references, and relationship to the study's actual evidence.

For each flagged sentence, complete a structured editorial assessment:

  1. Evidence: Is the claim supported by the study data, analysis, or a verifiable external source?

  2. Specificity: Does it describe this study, or could it appear unchanged in hundreds of unrelated papers?

  3. Provenance: Can the responsible author explain how it was drafted and edited?

  4. Attribution: Does it express an external finding, theory, method, or interpretation that requires citation?

  5. Classification: Does it look more like raw AI-generated text or AI-paraphrased writing?

  6. Disposition: Retain, substantiate, rewrite for precision, reframe and cite, or remove.

The review should generate notes, not merely changes. A rationale such as "Retained: human draft, supported by Table 2, regularized by grammar editing only" is stronger than an unexplained acceptance or deletion.

Step 3: Use sub-classification to choose the correct remedy

The same edit is not appropriate for every flagged sentence. A raw AI-style sentence often needs scientific specificity; an AI-paraphrased sentence needs source investigation and attribution review.

Risky raw-AI-style sentence

"The results demonstrate the urgent need for a comprehensive and interdisciplinary response to this growing problem."

The sentence may be fluent, but it is not yet a scholarly conclusion. It does not name the result, show the magnitude or direction of an effect, identify the study's limitations, or distinguish observation from recommendation.

Evidence-based revision

"In this sample of 214 participants, lower follow-up attendance was associated with treatment non-adherence after adjustment for age and baseline severity; prospective implementation studies should test whether reminder systems improve attendance in comparable settings."

The revision is not preferable because it is less "AI-like." It is preferable because it states an observable result, identifies an analytical boundary, and frames the recommendation as a future research question rather than an unsupported conclusion.

Now consider an attribution problem.

Risky AI-paraphrased sentence

"Stylometric systems use vocabulary and sentence patterns to distinguish writing produced by people from text created by artificial intelligence."

If this sentence reproduces the substance and structure of a source without citation, changes in wording do not solve the problem.

Attributed and compliant revision

"Stylometric approaches assess lexical and syntactic distributions to estimate whether a passage resembles human or AI-generated writing; because those estimates are probabilistic and context-sensitive, they should inform human review rather than serve as conclusive evidence of provenance (Author, Year)."

In a manuscript, the reference must point to an authentic, reviewed source. The revised text both attributes the underlying methodological idea and describes its evidentiary limit.

Step 4: Conduct DOI-linked source verification

For any sentence suspected of source dependence, use the similarity index to locate DOI-linked records. Read the relevant source context — not only the matching fragment. Compare wording, argumentative sequence, statistical details, definitions, and interpretive claims. Then record what you found.

If a citation exists but the manuscript closely reproduces the source's distinctive expression, add quotation marks where appropriate or rewrite more substantially. If no citation exists, decide whether the source is actually being used; if it is, cite and reframe it. If a match arises from the team's own thesis, preprint, protocol, conference paper, or prior publication, identify the prior record and follow the target journal's policy on disclosure, citation, permissions, and text reuse.

The purpose is not to drive the similarity index to the lowest possible value. The purpose is to ensure that every important match is intelligible, permissible, and documented.

Step 5: Review AI-use disclosure and authorship

After the text and source audit, compare the actual workflow with the journal's requirements. Determine whether use was limited to basic proofreading, involved substantive reorganization or generation, or extended into analysis, code, visualization, or data processing. Describe the use in the location and format the target journal specifies — often an acknowledgement, a declaration, the Methods section, and/or the cover letter.

A precise declaration is better than a vague one:

"The authors used [tool and version] to identify grammar and clarity revisions in selected portions of the manuscript. All suggestions were reviewed, edited, and approved by the authors, who take full responsibility for the accuracy, originality, and integrity of the submitted work."

If AI influenced the research process rather than writing alone, the methods description should state inputs, versions, validation steps, human oversight, and reproducibility details. Do not list an AI tool as an author or use a disclosure to excuse unverified output.

Step 6: Export a reproducible evidence record

When the audit is complete, download the highlighted PDF evidence report. Retain it with the final manuscript identifier, sentence-level disposition notes, DOI-linked source checks, AI-use log, final disclosure, and author approvals.

The report is not a certificate that the text is entirely human-authored or free from any integrity issue. Its value is evidentiary: it preserves what was identified, what was reviewed, what sources were consulted, and what actions the authors took. This allows a team to explain its process accurately if queried later.

Turn detection signals into editorial decisions. Start free on SciCampus — no card required — then export a highlighted PDF evidence report once each flagged sentence has a documented resolution. Run your first audit.

Edge cases and common pitfalls

False positives in non-native English and edited writing

A sentence with a red, orange, or yellow band is not proof that a non-native English-speaking author used AI improperly. Professional editing, translation, formulaic technical phrasing, and highly regular writing can all influence stylometric signals. The ethical response is not to introduce grammatical errors or artificial stylistic variation. Preserve drafts and editorial records, verify the claims, and ensure the author understands and endorses every sentence.

If an editor asks about a passage, explain the workflow factually: who drafted it, what editing was applied, and what evidence supports it. Provide documentation when requested. Avoid categorical claims that no detector can ever be useful; instead, request a contextual human evaluation of the text and evidence.

Multi-author workflow and shared accountability

In a large lab, one contributor may use AI for a section while another assumes it was independently drafted. This gap is avoidable. Designate a manuscript integrity lead — usually the corresponding author, a senior postdoc, or a project manager — who maintains the audit log, coordinates resolution of flagged text, and confirms that all authors review the final manuscript.

Agree in advance who runs each pre-submission audit and where the resulting reports are stored, so that review history, evidence reports, and source checks are retained consistently across the group. Access controls and data-classification rules should be established before any manuscript is uploaded to a third-party service.

Conventional methods language and legitimate reuse

Methods sections often include standard terminology, instrument descriptions, regulatory wording, or reporting language. Similarity and stylometry flags in such material require judgment. Do not assume all methods overlap is harmless, particularly when text is copied from prior work, a manufacturer's manual, or another team's protocol. Verify whether the language is standard, whether citation is needed, and whether the target journal expects a distinct description.

For text reused from an author's earlier output, distinguish acceptable continuity from redundant publication. Cite and disclose prior records when appropriate. Reusing a detailed results narrative, discussion, or conclusion is generally more problematic than using carefully attributed standard methodological language.

The detector-avoidance fallacy

No integrity workflow should have "make all colours disappear" as its objective. Attempts to evade a system — by obfuscating language, using multiple rewriters, introducing mistakes, or changing wording solely to alter a probability — can damage clarity and create a more concerning provenance record. Revise only for substantive reasons: unsupported claims, unclear methods, inaccurate citations, excessive closeness to a source, or lack of ownership.

Preparing for an editor or ethics inquiry

If a manuscript is questioned, prepare a concise, date-stamped case file rather than sending broad assurances. Include:

  • The exact manuscript version and the specific sentences queried.

  • Version history and tracked changes that show the drafting process.

  • A transparent AI-use log, including the tool, role, scope, and human review.

  • DOI-linked records for relevant matched sources and the team's disposition of each.

  • The highlighted PDF evidence report.

  • The final AI declaration, author-contribution record, and approval evidence.

This documentation does not ask an editor to trust a detector. It allows the editor or committee to conduct an informed review of the underlying evidence.

Verification checklist and lab SOP

Text and provenance review

  • Freeze and label the exact candidate submission version before beginning the final audit.

  • Confirm that each author has reviewed and approved the sections connected to their contribution.

  • Complete the AI-use log for substantive assistance: tool, version, date, purpose, section, human reviewer, and disposition.

  • Run a sentence-level review and examine every red (≥99%), orange (≥96%), and material yellow (≥93%) signal in context.

  • Use sub-classification to determine whether a flagged sentence appears to be raw AI-generated text, AI-paraphrased writing, or a possible false positive.

  • Verify each substantive claim against data, analysis outputs, or authentic cited evidence.

  • Check every citation and DOI directly; never rely on a generated reference without confirming the source record.

Similarity and attribution review

  • Use DOI-linked similarity evidence to investigate meaningful matches at sentence level.

  • Classify each match as standard language, quotation, properly cited paraphrase, legitimate self-overlap, or revision required.

  • Confirm that quotations are marked and cited, and that paraphrases are independent in both wording and argument structure.

  • Check overlap with theses, preprints, registered reports, protocols, conference outputs, prior articles, and grant materials.

  • Remove any text altered solely to evade AI or similarity detection rather than to improve scholarly accuracy and attribution.

Compliance and evidence retention

  • Read the current policies of the target journal regarding AI, authorship, confidential material, images, data, and declarations.

  • Draft a specific AI-use declaration when required, matching the actual workflow.

  • Confirm that no AI tool is listed as an author or contributor.

  • Download the highlighted PDF evidence report and preserve it with the final audit date.

  • Archive disposition notes, DOI-linked source checks, version history, AI-use logs, and author approvals with the submitted manuscript.

  • Confirm that the file submitted to the journal is the same version that completed the audit.

A five-gate lab SOP

A lab can standardize the process without making it burdensome:

  1. Initiation: Define permitted AI uses, privacy restrictions, contributor roles, and data-handling rules.

  2. Drafting: Preserve version history and document substantive AI assistance as it occurs.

  3. Audit: Conduct sentence-level stylometry and DOI-grounded similarity review; resolve each material finding.

  4. Compliance: Verify authorship, disclosure, citations, and target-journal instructions.

  5. Archive: Export the highlighted PDF evidence report and retain it with the approved submission version and supporting records.

Make your pre-submission review defensible. Audit your manuscript with SciCampus — free to start, no card required. Review sentence-level signals, trace similarity to DOI-linked sources, and retain an evidence record before you submit. Get started.

Frequently asked questions

What does a 30% AI score actually mean?

On its own, very little. The same percentage can describe a paper with two paragraphs of unreviewed chatbot output and a paper whose non-native English author used grammar assistance on entirely original work. The number tells you nothing about which sentences produced it, what they claim, or whether any change is warranted. Only sentence-level inspection answers that.

Is a high AI score evidence of misconduct?

No. Detection methods estimate whether text statistically resembles machine-generated writing. They cannot establish authorship or intent, and their output is affected by discipline, translation, copy-editing, document length, and model changes. A high score is a reason to review specific passages, not a finding against an author.

How is AI-paraphrased text different from raw AI output?

Raw output tends to be fluent but generic — broad claims disconnected from the study's data, sometimes with invented references. AI-paraphrased text is usually a rewording of an existing source that preserves its conceptual structure. The first needs scientific specificity; the second needs source verification and correct attribution. Treating them the same way leads to the wrong fix.

Should I rewrite sentences until the flags disappear?

No. Editing solely to change a probability score can degrade methodological precision and, if it conceals source dependence, create a worse provenance record than the original. Revise when a sentence is inaccurate, generic, unsupported, improperly attributed, or not genuinely yours.

What should a non-native English speaker do about false positives?

Keep evidence rather than change style. Preserve drafts, tracked changes, editing correspondence, and source notes, and make sure you can explain and defend every sentence. If an editor raises a passage, answer factually with that documentation and request a contextual human review.

What belongs in a pre-submission evidence record?

The frozen submission version, version history, an AI-use log covering tool and purpose and human reviewer, DOI-linked source checks with a disposition for each match, the highlighted PDF report, and the final disclosure statement with author approvals. Together these show that decisions were made deliberately and can be explained months later.

Conclusion

Single aggregate AI scores fail because they summarize uncertainty without preserving the information needed for scholarly judgment. They cannot identify the sentence at issue, discriminate between raw AI output and AI-paraphrased writing, establish attribution, resolve a false positive, or demonstrate that authors exercised responsible oversight.

A defensible pre-submission audit works at the level where academic integrity is actually assessed: the sentence, the claim, the source, the draft, and the documented author decision. Sentence-level stylometry provides a targeted review queue. Sub-classification directs authors toward the right remedy. DOI-linked similarity checks make source review concrete. Highlighted PDF reports help research groups retain a coherent evidence record across collaborative projects.

The purpose is not automated clearance. It is better scholarship: precise claims, verified sources, transparent methods, accountable authorship, and a manuscript that can withstand informed editorial scrutiny.

Policy note: Publisher and journal requirements change. Always consult the current author instructions for the exact target journal; where those instructions are more restrictive than general guidance, the journal's policy governs.

Policy resources

Comments

No comments yet — be the first to share your thoughts.

Leave a comment

Comments are reviewed before they appear. Your email is never published.

Put this into practice

Start free — detect AI content, match journals, and get more done with SciCampus.

Keep reading

All blog