Linda-Stylo-Clean 1.0
Stylometric AI-text detector for English (~250 style features + hashed n-grams -> LightGBM). CPU only, milliseconds per text, no GPU, no cloud. Trained only on data that pass a strict clean-provenance filter (open/permitted sources; AI texts generated by open-weight models with permissive terms). See DATA_BOM.md.
License: free for research / non-commercial use (CC BY-NC 4.0). Commercial use is paid — see COMMERCIAL_LICENSE.md, contact lindapro.support@proton.me.
pip install -r requirements.txt
python -m linda_stylo your_text.txt
from linda_stylo import LindaStylo
print(LindaStylo().detect(["text ..."])) # score + verdict (ai above the 1%-false-positive threshold, uncertain above 5%)
Measured (Chicago Booth benchmark, 1% threshold on the benchmark's humans): plain AI 97.7%, after StealthGPT 54.9%; false flags on school ELL essays 2.5%, TOEFL 2.2%. It is weaker than Linda-Pro, which adds two transformer voters, and it misses polished essays by modern models. Use it when data lineage, CPU-only operation or speed matter more than peak accuracy. English only. Not the sole basis for decisions about people. Contact: lindapro.support@proton.me