← projects

AI-Generated Text Classifier

ML research project — Classical models with engineered readability features that beat a fine-tuned BERT baseline.

Python · scikit-learn · BERT · NLP · SHAP · Jupyter

Where
CS 463/663, Foundations of Machine Learning
Team
Coauthored with Ezequiel Beck

Problem

The default answer to AI-text detection is to fine-tune a transformer, which is expensive to train, expensive to run, and hard to interrogate when it is wrong. It was not obvious that the expense was buying anything.

Solution

Train Logistic Regression and Random Forest on TF-IDF vectors combined with engineered readability signals — Flesch, Gunning Fog, SMOG, length, word count — and put them head to head with a fine-tuned BERT on the same task and split, then use SHAP to see which features actually carried the decision.

Overview

A study in whether interpretable, low-resource models can hold their own against a transformer on AI-text detection. Logistic Regression, Random Forest, and a fine-tuned BERT were trained on the same task — labeling text as human-written or AI-generated — using TF-IDF vectors capped at 10,000 features combined with engineered structural and readability signals: Flesch Reading Ease, Gunning Fog, SMOG index, character length, and word count. Data came from a labeled Kaggle essay set plus the artem9k/ai-text-detection-pile corpus of over 1.3 million examples, split 80/20 with grid cross-validation and interpreted using SHAP. Random Forest reached 99% accuracy and Logistic Regression 94%, while the BERT baseline reached only 80% — the engineered readability features turned out to carry most of the signal. Coauthored with Ezequiel Beck for CS 463/663, Foundations of Machine Learning.

Impact

  • Random Forest 99% accuracy, Logistic Regression 94%, fine-tuned BERT 80%.
  • SHAP attribution showed the engineered readability features carried most of the signal.
  • Trained across a labeled Kaggle essay set plus a 1.3M-example corpus, 80/20 split with grid cross-validation.