Hate Speech-Detecting AIs Can Be Fooled: Study

Advertisement
By Indo-Asian News Service | Updated: 17 September 2018 17:01 IST

Machine learning detectors deployed by major social media and online platforms to track hate speech are "brittle and easy to deceive", a study claims.

The study, led by researchers from the Aalto University in Finland, found that bad grammar and awkward spelling - intentional or not - might make toxic social media comments harder for artificial intelligence (AI) detectors to spot.

Advertisement

Modern natural language processing techniques (NLP) can classify text based on individual characters, words or sentences. When faced with textual data that differs from that used in their training, they begin to fumble, the researchers said.

"We inserted typos, changed word boundaries or added neutral words to the original hate speech. Removing spaces between words was the most powerful attack, and a combination of these methods was effective even against Google's comment-ranking system Perspective," said Tommi Grondahl, a doctoral student at the varsity.

Advertisement

The team put seven state-of-the-art hate speech detectors to the test for the study. All of them failed.

Among them was Google's Perspective. It ranks the "toxicity" of comments using text analysis methods.

Advertisement

Earlier, it was found that "Perspective" can be fooled by introducing simple typos.

But, Grondahl's team discovered that although "Perspective" has since become resilient to simple typos, it can still be fooled by other modifications such as removing spaces or adding innocuous words like "love".

Advertisement

A sentence like "I hate you" slipped through the sieve and became non-hateful when modified into "Ihateyou love".

Hate speech is subjective and context-specific, which renders text analysis techniques insufficient as stand-alone solutions the researchers said.

They recommend that more attention be paid to the quality of data sets used to train machine learning models - rather than refining the model design.

The results will be presented at the forthcoming ACM AISec workshop in Toronto.

 

Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.

Further reading: AI
Advertisement

Related Stories

Popular Mobile Brands
  1. GTA 6: An Extended Look Debuts on Netflix Today: How to Watch, What to Expect
  2. Crimson Desert's New 'Enhanced' Update Attemps to Fix the Game's Story
  3. These Are the Best One-Time Investment Laser Printers in India
  4. Realme C100i With a 6,500mAh Battery Debuts in India at This Price
  5. OnePlus 16 Display Tipped to Offer Ultra-Thin Four-Sided Bezels
  6. Realme P4 Power, P4 Lite Prices Hiked in India by Up to Rs. 3,000
  1. Vivo X500 Pro Max Display Details Officially Teased Ahead of Launch
  2. Apple Maps Gets Transit Support in India With Public Transport Directions and More
  3. Google Is Changing How You Sign In to Android Apps
  4. Sneaky Sasquatch Gets Subway Surfers+ Crossover, Birdwatching Update on Apple Arcade
  5. YouTube Adds Amazon Product Tags to Shorts, Videos and Livestreams
  6. Poco X8 5G and X8 Power 5G India Launch Date Set for September 4
  7. Acer Swift Neo 14 Refreshed in India With Up to Intel Core Ultra 7 CPU: Price, Features
  8. Samsung Odyssey G Series Gaming Monitors Unveiled at Gamescom 2026 With Curved OLED, 1,100Hz Panel
  9. Google's AI Mode Can Now Track Flight Prices and Book Hotels for You
  10. Google Unveils Gemini Omni 1.1 Flash With 4K Video and 40-Second Extension Support
Download Our Apps
Available in Hindi
© Copyright Red Pixels Ventures Limited 2026. All rights reserved.