Structural and Linguistic Markers for Distinguishing Human Manuscripts from AI-Generated Summaries


Pit Pichappan
Digital Information Research Labs Chennai 60017, Tamil Nadu. India
pichappan@dirf.org



ABSTRACT

The rapid development of Large Language Models (LLMs) has made it increasingly hard to distinguish human written from AI-generated academic papers, raising questions about authorship, originality, and academic integrity. This research presents a systematic approach to identifying and measuring multidimensional markers that differentiate human written academic manuscripts from AI-generated and AI-paraphrased summaries. After analysing 52 human written academic papers and their AI-generated versions (created with GPT-4o, DeepSeek V4, and Qwen 3.7plus), we developed an automated, multi-layered text-annotation tool. It incorporates rule based matching, zero shot classification, and structured LLMaided annotation, and we validated it using human inter annotator agreement. We found that no single language marker is enough for classification. However, the accurate classification (85-90%) requires a set of markers which include the completeness of the scholarly apparatus (the most reliable one), significant information loss (40-50% of granular details such as exact statistics or software name), decrease of syntactic burstiness (by 40-50%), formulaic transition substitutions, and the drop in epistemic hedging by 60%. Although the absence of structure and metadata yields perfect classification, the stylometric approach alone produced an AUC of 0.78, underscoring the complexity of linguistic signatures. In summary, successfully implementing AI for identifying text in academic environments requires a multi marker approach that favours completeness, informativeness, and text depth over fluency and vocabulary sophistication, which are highly prone to false positive results. Previous research has addressed AI academic detection; however, we shift our emphasis from surface fluency (which LLMs excel at mimicking) to epistemic rigour and granular information loss.

Keywords:

AI-generated Text Detection, Large Language Models (LLMs), Academic Writing, Stylometric Analysis, Information Retention, Scholarly Apparatus, Automated Text Annotation, Syntactic Burstiness, Epistemic Hedging, Multi-marker Evaluation


ALL METRICS

185

VIEWS


158

DOWNLOADS


Get PDF Get XML Cite Export
Introduction

Generative AI (LLMs) reflects one of the most transformative advancements in Natural Language Processing. These models are trained on massive corpora of text and other data and have shown an unprecedented ability to understand, generate, and manipulate natural text. Their applications span almost every sector and discipline, and users increasingly rely on them. At the same time, users tend to depend on them too heavily for text generation, challenging human capacity to generate knowledge.

Related Work

The speedy growth of generative artificial intelligence (AI), particularly Large Language Models (LLMs), has greatly changed the ways in which educational and professional realms create and apply written information. While web tools and electronic resources were already central to academic communication, generative AI amplifies their use and introduces new ethical and intellectual issues. More and more, authors use machinewritten and machine assisted texts, and hybrid texts that mix human generated and machine generated content are increasingly common. This new situation has raised questions about authorship, originality of ideas, and the role of machine writing in academic writing. Modern LLMs can produce logical, meaningful text, making it difficult to distinguish between human written and AI-generated texts [1, 2]. However, research shows that AI-generated texts can contain serious factual mistakes, underscoring the need to use generative AI responsibly and systematically evaluate AI-generated texts [3].

The popularity of LLMs relies heavily on advances in computational linguistics and natural language technology [4]. Systems such as ChatGPT also create logical, context-relevant texts and raise many questions concerning authorship and textuality [5]). Researchers have done significant work identifying linguistic and stylistic features that might help differentiate AI-generated text from human writing. However, literature increasingly points out the impossibility of establishing this distinction based on any one linguistic feature or universal marker. A systematic review of differences between AI-generated and human writing showed that this distinction relies on multilayered, situational profiles of signals rather than a single effective indicator [6]. In addition, some studies assessing current LLM text-detection methods argue that multilinguality, hybrid texts, shifting language generation models, and regulation remain long term obstacles to reliable detection [7].
The limitations detected further complicate identifying machine generated text. Studies show poor and inconsistent results from existing detection tools [8-12]. These limitations matter because human evaluators also cannot distinguish machine written work from human written work.
AI-detection tools can flag suspicious texts, but their results should not be treated as concrete proof of improper AI use. Human judgments are often highly diverse and inconsistent [13]. The advancement of LLMs makes it even harder for even the most competent readers to determine whether a text was written by a human or an AI [14].

2.1 Linguistic Characteristics of AI-Generated and Human-Written Text

One central method for differentiating AI- and human-generated text is to analyse it at the lexical, syntactic, morphological, and readability levels. Studies show that AI-generated texts differ systematically from humanwritten essays. Studies report that AI-generated writing shows lower lexical diversity, greater syntactic complexity, and higher nominalisation than human writing [15, 16, 17]. This suggests that AI-generated texts should have a distinct linguistic profile, even when their content is coherent and grammatically correct. Fedoriv’s [18] research demonstrates that potential markers of AI-generated English texts might include repetition, illogicality, tone mismatches, and factual mistakes. However, these features should not be treated as definitive indicators of AI authorship, since language-generating technologies are still evolving.
The illustrations suggest differences in vocabulary use and structural organisation. Texts produced by GPT have be- en noted to use a smaller vocabulary, have little lexical variation, and repeat phrases frequently. In addition, the literature assumes that GPT-produced texts show specific characteristics in terms of syntactic complexity, morphology, and readability [19]. This supports the suggestion that GPT texts may have a sort of “register” characterised by specific recurrent structures. However, these features may differ across domains.
Comparative statistical analyses of previously mentioned texts provide more evidence. Human writing has less complex syntax but more semantic diversity, whereas human written texts show significant stylistic variation. At the same time, human texts show greater variability in the discussed linguistic features [20]. The results highlight the need to analyse multiple linguistic dimensions at once, rather than individual features.
Morphosyntactic features can also help distinguish the two. Specifically, AI-generated texts are believed to be more nominally loaded, with a higher density of nouns and less frequent use of pronouns and auxiliary verbs. Human texts are generally more morphologically complex when it comes to their reference; tense; and aspect [21, 22, 23, 24]. This means AI and human texts may differ not only in vocabulary and grammar but also in how authors express their viewpoints and references.

2.2 Stylistic, Structural, and Discourse-Level Characteristics

Beyond individual words, writing organisation and flow also differ between human and AI writing. Studies in different academic fields show that these organisational and flow features can distinguish student writing from AI-generated text [25]. This is especially true in school settings where writing needs to show more than just correct grammar. It should also show how well someone understands the subject, can build an argument, take a stand, and understand the context.
Some research has found that AI writing often has certain structural patterns. These include disconnected arguments, predictable organisation, specific language patterns, and an overly objective tone. [26]. This is different from how humans often write, which tends to be more connected and subtle. Also, whether people think a text is AI-generated isn’t just about grammar. Teachers have pointed out that a lack of language mistakes, complicated sentence structures, advanced vocabulary, made up information, fake sources, repeating words or phrases, and noticeable text structures can all be signs of AI writing. [27]. This suggests that our judgment of AI text comes from a mix of factors: grammar, word choice, organisation, factual accuracy, and stylistic consistency.
Some research has found that AI writing often has certain structural patterns. These include disconnected arguments, predictable organisation, specific language patterns, and an overly objective tone. [26]. This is different from how humans often write, which tends to be more connected and subtle. Also, whether people think a text is AI-generated isn’t just about grammar. Teachers have pointed out that a lack of language mistakes, complicated sentence structures, advanced vocabulary, made up information, fake sources, repeating words or phrases, and noticeable text structures can all be signs of AI writing. [27]. This suggests that our judgment of AI text comes from a mix of factors: grammar, word choice, organisation, factual accuracy, and stylistic consistency.
It’s getting harder to tell the difference because AI can now produce content that fits the context. Even though these systems can create text that sounds fluent and appropriate, worries remain about whether the information is accurate and whether they can capture the complex aspects of authorship typical of human writing. [28] These concerns go beyond surface level language differences, and touch on intent, point of view, originality, and how language relates to human thought. Because of this, some suggest using psycholinguistic approaches to find patterns in human writing and develop better ways to ensure academic honesty as AI becomes more common. [29].

2.3 Development of AI-Generated Text and Detection Approaches

The lines between human and AI writing are blurring because AI text generators keep getting better. As these AI systems are updated, they change how they write and what they produce, making their output look more and more like human writing. This means that tools designed to detect AI text need to keep up with these changes [30]. Some research shows that AI-generated essays can be longer, use more complex sentences, use a wider vocabulary, and include more information than human-written essays. Even with these differences, AI and human writing can express feelings and connect ideas logically in similar ways. AI writing might even use more complex sentence structures and advanced ways of building sentences [30]. This means telling AI writing apart from human writing isn’t a simple, fixed task. The old idea that AI text is always repetitive, simple, or basic is probably no longer true as AI gets smarter. Things that used to be clear signs of AI writing might be gone or harder to spot in newer AI versions. At the same time, new patterns might emerge as AI gets better at sounding human. So, to reliably identify AI text, we need a flexible, constantly updated approach.
Scientists are also exploring machine learning to better distinguish AI writing from human writing. For instance, one researcher suggested combining a genetic algorithm and a multilayer perceptron with a BERT based method. Both performed better than older detection methods [31]. This shows a shift away from simple, set detection rules toward computer based methods that consider many aspects of the text. However, how well these methods work depends on the variety of data they are trained on, the specific AI language models used, and whether the testing data actually includes examples of current writing that mixes human and AI input.

2.4 Performance of the detection tools

AI detection tools’ reliance on standardised linguistic metrics, such as perplexity and burstiness, has raised significant concerns about bias against non-native English speakers (NNES) [32]. For instance, Liang et al.[33], found that GPT-based detection tools misclassified over half of the text samples from NNES, yielding an average false positive rate of 61.3%. However, industry research contests this finding; studies associated with GPTZero and Turnitin claim their tools can accurately identify human written text from NNES without showing classification disparities [34, 35].
Despite these claims, broader concerns about the reliability of AI detectors remain. For example, OpenAI’s AI classifier was criticised for not including NNES generated text in its training data, even though it was trained on varied human textual patterns [36, 37]. Because it could not reliably detect generative AI (GenAI) output, OpenAI withdrew the classifier in July 2023. In fact, research frequently contradicts overarching claims of detector accuracy by showing their highly variable and often poor ability to distinguish AI and human generated content [38-44].
A major factor contributing to this unreliability is detectors’ vulnerability to text manipulation. Studies have shown that detection accuracy plummets when GenAI output is altered. Mitchell et al. [45] (2023) reported that while DetectGPT initially identified 70.3% of GPT2 XL generated sequences correctly, this rate fell to just 4.6% after the content was processed through an automated paraphrasing tool (APT). Similarly, Weber-Wulff et al. [44] observed major drops in accuracy after using both APTs and translation tools.
These vulnerabilities highlight the inherent limitations of current AI text detectors. As Originality.AI [46] notes, detectors consistently lag behind recent GenAI advancements. They struggle to distinguish increasingly complex AI-generated content from human writing and remain highly susceptible to adversarial bypass techniques. Consequently, these compounding problems underscore the urgent need for educational institutions to balance AI detection tools with policies that accommodate AI-produced materials [47].

<
2.5 Research Gap

Even as we learn more about AI-generated text, the research remains scattered. It is split across areas like language, style, computer science, and methods for spotting AI text. Many studies have examined specific traits of AI writing, but they haven’t synthesised these findings into a clear picture of how AI and human writing differ across multiple dimensions.
Previous reviews [48, 49] mainly focused on the broader task of detecting AI text, or on whether large language models (LLMs) understand their own language limits. As a result, little work has systematically brought together the general language patterns [50] we see when comparing AI and human writing. [51-54]. So, the research shows we need ways to look at more than mere one angle. We need to consider how word choice, sentence structure, word formation, meaning, style, overall structure, cohesion, and conversational language use work together. This is especially true as LLMs advance, AI writing varies widely by subject, people start writing with AI, and newer models keep changing the language they produce. By systematically examining these aspects, we can better understand how the differences between human and AI writing are changing. This will also give us a stronger foundation for creating and testing tools that automatically label and detect AI-generated text.

3. Methodology

We used 52 independent human written academic texts (452 MB) spanning addiction science, complex systems, digital health, SUD interventions, clinical care, scientific editing, biomedical methods, and soil/irrigation science. The collection spans diverse disciplines, authors, styles, and lengths, and includes both published and unpublished work.
Model A (Full Model): The current model using all 23 features (including credit_section_present, citation_density_per_1k, specific_software_named).
Model B (Stylometric-Only Model): Retrain the Gradient Boosting classifier using only linguistic, semantic, and stylistic features (e.g., transition variety, sentence length variability, hedging frequency, synonym substitution rate,burstiness).
We generated AI counterparts using GPT-4o, DeepSeek V4, and Qwen 3.7plus under controlled conditions (temperature 0.7, top-p 1.0, max tokens 8,192, seed 42). A standardized system prompt instructed the models to act as expert academic editors and rewrite each source into a clear, concise, well-structured academic summary while preserving core meaning, findings, methods, and conclusions.
We generated AI counterparts using GPT-4o, DeepSeek V4, and Qwen 3.7plus under controlled conditions (temperature 0.7, top-p 1.0, max tokens 8,192, seed 42). A standardized system prompt instructed the models to act as expert academic editors and rewrite each source into a clear, concise, well-structured academic summary while preserving core meaning, findings, methods, and conclusions.
Generation strategies include the following:

  • Paraphrasing – full source text provided for rephrasing and condensation
  • Direct generation – writing from instructions without the source
  • Hybrid – partial human editing followed by AI regeneration

3.1 Research Design and Automated Annotation Framework

This study uses a structured automated text-annotation framework to systematically describe and distinguish between human generated and AI-generated text. The method turns the research goal into a repeatable annotation scheme. This scheme includes set labels, structural rules, feature definitions, and decision criteria. The framework aims to reduce subjective interpretation while ensuring the annotation process is consistent, transparent, and repeatable. The proposed framework combines automated annotation methods that use pretrained language models, zero-shot classification, rule-based linguistic matching, and large language models (LLMs). Rather than relying on a single detection method, this approach combines multiple methods to identify text features at the meaning, structure, language, and discourse levels. This multi-faceted strategy aligns with research showing that no single language marker can reliably distinguish AI-generated text from humangenerated text.

3.2 Development of the Annotation Schema

The first step was to create an annotation schema that translates the research objectives into a formal set of labels and rules. The schema defines the categories to be identified, the characteristics linked to each category, and the boundaries for assigning them. We paid special attention to creating clear, unambiguous annotation criteria to reduce interpretive variation. The annotation framework was designed to handle different kinds of text information. This includes structural markers, meaning based characteristics, language features, and indicators related to how synthetic or machine generated the writing is. The resulting classification system provides a shared representation that allows different annotation tools to assign labels consistently across texts.

3.3 Automated Annotation Approaches

We used automated annotation with complementary computational methods. Pretrained models, zero-shot learning, rule-based systems, and LLM-based annotation were used to avoid extensive manual labelling. These methods use existing language knowledge, pretrained neural representations, or clearly defined rules to find and classify text features. For semantic annotation, Hugging Face zero shot classification pipelines can assign semantic categories to text segments. Unlike traditional supervised classification, zero shot classification
allows you to use candidate labels without a task-specific training dataset that has been manually labeled. This means you can evaluate text against a set of candidate categories, and the model assigns the most likely semantic label. For identifying structural and text markers, rule based methods can find clearly defined patterns. Specifically, spaCy with regular expressions or rule based matching can identify predefined text markers. GLiNER can perform zero shot named entity recognition when the target categories are not limited to a standard, fixed set of entities. These methods are good for automatically finding structural expressions and markers within documents.

3.4 LLM-Based Structured Annotation

LLMs were included as an extra annotation method for characteristics that need interpretation based on context and discourse. Local LLM environments, such as Ollama, vLLM, and LM Studio, process sensitive text without sending data to external API services. Models like Llama-3 or Qwen can be run locally and integrated with Python based batch processing workflows for large scal e automated annotation. We used a structured output approach to make sure LLM generated annotations fit predefined fields instead of producing openended text responses. This lets us evaluate each text unit against a consistent annotation classification system and return it in a machine readable format. Structured outputs also make it easier to validate later, filter by confidence, check for consistency, and build the final annotated dataset.

3.5 Assessing AI-like Qualities

We can use a language model to evaluate how much a piece of text sounds AI-written. This involves giving the text a “synthetic score.” We can check individual pieces of text or compare pairs of texts based on things like how well the ideas connect, how transitions are used, and how information is woven in. The results are given as structured data, not merely a general opinion. This way, the AI-identified synthetic traits remain specific annotation fields that can be checked for accuracy and consistency. We remember that this score is only a measure, not proof that AI wrote it, because AI text detection is not perfect, according to research.

3.6 Identifying Themes and Language Features

Another part of this process is finding the themes and features within the text. We use tools like Hugging Face’s zero-shot classification and SciSpaCy to automatically sort paragraphs into categories. For example, categories could include method adaptability, applicable rules, or social and environmental factors. This process turns a collection of plain text into a structured format with different categories and their related features. It makes it easier to compare human and AI written texts and to examine language, meaning, and themes together rather than separately.

3.7 Putting It All Together

The whole annotation process is a step by step workflow. It starts with the raw text, then generates labels by category, validates the structured results, filters out low confidence findings, checks for consistency, and finally gets independent validation. The end result is a clean, structured dataset ready for analysis. Using several annotation methods gives us different kinds of information and means we do not have to rely on one tool or set of rules. Rule based methods work well for clear structural markers. Pretrained and zero shot models can label large amounts of text efficiently for meaning. Language models help interpret more complex text features based on context.

3.8 Checking for Accuracy and Quality

Because automated annotation can make mistakes and introduce bias, checking accuracy is crucial. We validate the automated annotations by checking their structure and consistency before adding them to the final dataset. Filtering by confidence level helps us identify annotations where the model was not very sure. Independent validation adds another layer of quality control. A mixed approach that combines language models, rules, pretrained models, and some human review is best when accuracy and repeatability are the keys. The entire process, including the categories used, instructions given, tools and settings used, confidence levels, and how we checked things, needs to be documented. This makes the process clear, repeatable, and open to external review.

3.9 Making It Repeatable and Clear

To make sure others can repeat this work, we clearly state everything involved in the annotation process. This model includes the categories, their definitions, the software used, model settings, instructions, rules, confidence levels, and validation methods. Automated annotation should use the same settings throughout the text collection, and the structured results should be kept with the data for analysis. The resulting methodology provides a reproducible framework for converting heterogeneous textual material into structured annotations while combining semantic, linguistic, structural, and synthetic content indicators. It is consequently suitable for systematic comparison of human generated and AI-generated text and for subsequent development or evaluation of automated text annotation and AI-generated content detection approaches.
This analysis consolidates findings from thirty two comparative analyses of human written academic text and AI-generated or AI-paraphrased academic text. The evidence is organised around lexical, semantic, stylistic and syntactic, formatting and typographical, provenance and metadata, scholarly apparatus, and annotationrelated features. The central finding is that no single linguistic feature is sufficient for dependable classification. Stronger differentiation emerges when multiple independent markers are considered together, particularly academic scaffolding, information specificity, sentence variability, transition patterns, and the depth of critical analysis.

3.10 Inter-Annotator Reliability Assessment

To ensure the validity and consistency of our automated annotation framework, we conducted a systematic inter annotator reliability assessment. While the primary annotation pipeline relies on automated methods (zero shot classification, rule-based matching, and LLM based annotation), human validation remains essential for quality control and for establishing benchmark reliability against which automated performance can be compared.

3.10.1 Annotator Selection and Training

A panel of three independent annotators participated in the reliability assessment. Annotators independently evaluated a stratified random sample of texts drawn from the full dataset:
Annotators evaluated each text on the following categories, using the same schema as the automated framework: Transition Quality, Lexical Diversity, Syntactic Complexity, Information Density, Scholarly Apparatus Completeness, and overall AI-likeness and Confidence Rating.

Component Details
Sample size 104 texts (52 human-written, 52 AI-generated)
Stratification Balanced across disciplines, text lengths, and AI models
Annotation units Full documents (for holistic assessment) plus 50 paragraph level units (for fine grained analysis)
Blinding All identifiers removed; annotators were not informed of the text source (human vs AI)
4. Results

4.1 Lexical Markers
4. 1.1 Transition Words and Phrases
One noticeable difference is how transition words are used. The study found that the phrase “In contrast” in human writing was consistently replaced with “By contrast” in AI-generated text. This kind of consistent change suggests the AI is mechanically rephrasing, not naturally varying its words. People writing academically tend to use transition words more naturally and in different ways depending on the situation. AI writing, in contrast, often sticks to a smaller, more repetitive set of standard transitions. The comparison also showed that AI text uses phrases like “Furthermore,” “Nevertheless,” “Consequently,” “Importantly,” “This type of,” “More broadly,” “These developments suggest...,” and “This substantial discrepancy indicates...” more often. Words like “Additionally” and “Importantly” can appear in both human and AI writing, so just seeing them does not indicate much. More telling is whether they appear repeatedly or in a formulaic way, especially if the same transition pattern repeats throughout a document. Human writing usually has a wider variety of transitions and repeats them less predictably.