ShieldFont Poisons AI Training Data with Subtle Text Changes
TL;DR. Designers Isaque Seneda and Gabriel Abrucio created ShieldFont, a new font to disrupt AI model training by serving altered data to scrapers. - The font uses ligatures to replace words with semantically incorrect but grammatically similar substitutions for bots. - Human readers see the original, readable text, while scrapers collect corrupted data, impacting model quality. - This technique aims to provide web publishers with an opt-out from unauthorized data scraping and pollution.
- ShieldFont utilizes ligatures to substitute words, creating nonsensical content for AI scrapers.
- The font ensures that web pages remain perfectly readable for human users, masking the data poisoning.
- The primary goal is to disrupt unauthorized AI training by injecting misleading data into datasets.
- The method targets the semantic value of scraped text, making it useless for language model training.
Sources
- The web’s newest weapon against AI scrapers is a font — arstechnica.com