• The Closet Blog
  • Digital Collage and Multi Media Art
  • Excavation Extraordinaire
  • About the Artist
Menu

James Behan

  • The Closet Blog
  • Digital Collage and Multi Media Art
  • Excavation Extraordinaire
  • About the Artist

Digital photocollage and multimedia works documenting queer identity, gay male visibility, and the lives of men forced to live in the shadows. For Adult audiences. This site contains artwork depicting the gay male figure including nudity. Viewer discretion is advised.

Whose Content Survives: AI Moderation and Who Gets Silenced

August 10, 2026

The baseline that received more lenient treatment in every case was standard English, male-coded, white, and conventionally gendered. We all know who that is, even if AI will not admit it.

When people post on social media, most assume their content will be treated the same as anyone else’s — that the same words, images, and topics get the same treatment regardless of who posted them. The documented record shows that assumption is false. AI content moderation systems flag, remove, demonetize, and shadowban content from people of color, women, and LGBTQ+ users at higher rates than comparable content from other users, for reasons tied to who they are rather than what they posted.

People of Color

AI hate-speech detectors misread African American English (AAE) as hostile. A 2019 University of Washington study tested several detection tools, including Google’s Perspective, against 5.4 million tweets where the authors’ race was known. The tools were one-and-a-half to twice as likely to flag posts by users who identified as African American as toxic. [1] The cause was traced to the training data itself: human annotators had disproportionately labeled AAE tweets as offensive, and that bias carried into the algorithms trained on their labels. [1]

Later research confirms this wasn’t a one-off. A Carnegie Mellon study found a high correlation between annotators’ perceptions of toxicity and markers of AAE, causing AAE text to be mislabeled as abusive at a high false-positive rate. [2] A more recent benchmark of a widely used toxicity model found it scored AAE text as 1.8 times more toxic on average, and 8.8 times higher specifically for “identity hate.” [3] The practical result is content removal, shadowbanning, and demonetization of Black creators for ordinary language.

Women

AI moderation systems over-flag women’s bodies and women’s health content. Advocacy groups report that educational posts about breastfeeding, menopause, fertility, or sexual assault are removed or restricted, while comparable posts using male anatomical terms remain visible. [4] Clinical terms like “vagina,” “breast,” “menstruation,” or “cervix” are frequently misread as sexual content rather than health information. [4]

The clearest documented case is Instagram’s repeated removal of a photo of Nyome Nicholas-Williams, a plus-sized Black woman, covering her breasts with her arms — while comparable images of thin white women stayed up. The case became a public campaign (#IWantToSeeNyome), and researchers who studied it concluded the pattern reflected what they called a “pervasive platform policing of the female body,” with platform guidelines effectively inviting users to flag other women’s bodies as violations. [5]

LGBTQ+ Users

GLAAD’s Social Media Safety Index has repeatedly found platforms under-enforcing against anti-LGBTQ hate speech while over-enforcing against ordinary LGBTQ content — removing, demonetizing, and shadowbanning it without meaningful transparency. [6] YouTube creators have reported videos demonetized simply for containing “gay” or “lesbian” in the title. [7] Instagram has been found to police transgender-body content more aggressively than comparable cisgender content. [8] The Electronic Frontier Foundation and GLAAD have documented AI moderation systems flagging LGBTQ+ content as sexually explicit when no explicit material was present at all. [9] A 2024 USC study found posts using non-binary or gender-nonconforming language were flagged as “toxic” significantly more often than comparable posts using conventional gendered language, even when the content was harmless. [10]

Who Isn’t Being Flagged

If certain groups are over-moderated, the comparison built into these same studies shows who isn’t. Each study above measured the flagged group against a baseline that was treated more leniently by the same system, using the same criteria:

• Standard American English was the baseline against which AAE was found 1.5 to 8.8 times more likely to be flagged. [1] [3]

• Male-coded anatomical language was the baseline against which women’s health terms were found more likely to be flagged as sexual content. [4]

• Images of thin white women were the baseline against which the photo of a plus-sized Black woman was repeatedly removed. [5]

• Conventional gendered language was the baseline against which non-binary language was found more likely to be flagged as toxic. [10]

References

1. Futurism / Forbes, reporting on University of Washington study by Maarten Sap et al., “Google’s hate speech-detecting AI is biased against black people,” 2019.

2. Xia, Field, and Tsvetkov (Carnegie Mellon University), “Demoting Racial Bias in Hate Speech Detection.”

3. “How AI Fails: An Interactive Pedagogical Tool for Demonstrating Dialectal Bias in Automated Toxicity Models,” arXiv, 2026.

4. Healthline, “The Shadow Banning of Women’s Health Content.”

5. Frontiers in Communication, “Algorithmic agency and ‘fighting back’ against discriminatory Instagram content moderation: #IWantToSeeNyome,” 2024.

6. GLAAD, “Fourth Annual Social Media Safety Index.”

7. ArXiv, “Examining Multimodal Gender and Content Bias in ChatGPT-4o.”

8. Tech Policy Press, “The Censorship of LGBTQ+ Content Online Corresponds with Declines in Freedom for Everyone.”

9. UCLA eScholarship, “Algorithmic Bias in LGBTQ+ Content Moderation.”

10. USC Viterbi School of Engineering, “Flagged for Being Queer,” 2024.

Gays on the Titanic! →

Collage Blog Find

Velacirapting

Powered by Squarespace