Exploring the Impact of Linguistic Features on Pre-trained Models for Machine-generated Text Detection in Spanish

Javier Alonso Villegas Luis, Marco Antonio Sobrevilla Cabezudo

Resumen


Detecting machine-generated Spanish text remains challenging across domains and generators. Pre-trained models like RoBERTa provide strong contextual embeddings but often underperform on human-authored texts and are sensitive to domain shifts. In this work, we integrate linguistic features from PUCPMetrix—covering lexical, syntactic, semantic, psycholinguistic, and cohesion properties—with pre-trained models. We evaluate feature-based classifiers, fine-tuned RoBERTa, hybrid models, and ensembles on the AuTexTification dataset. Hybrid models improve human-text detection (F1 65.49 vs. 60.74 for RoBERTa) and machine-text classification (F1 81.76), while a voting ensemble achieves the highest macro-F1 (74.75) and strongest robustness. Analyses indicate linguistic features provide stable, interpretable anchors, reducing overfitting and enhancing generalization across LLM outputs. Results demonstrate that combining linguistic and pre-trained models yields a robust solution for Spanish machine-generated text detection.

Texto completo:

PDF