Skip to main content
mdfahd
Gainsight Employee ⭐️⭐️
September 10, 2025

How to Use the Moderation AI Agent

  • September 10, 2025
  • 44 replies
  • 1505 views

 This article helps moderators understand how to use Gainsight’s AI Moderation.

 

Overview

 

Community moderators are responsible for ensuring that user-generated content aligns with established guidelines and maintains a respectful, trustworthy environment. However, manually reviewing high volumes of content can be time-consuming, delay content publishing, and introduce inconsistencies in moderation.

Using the AI Moderation from the Pre-Moderation Rules in the community settings, moderators can streamline this process. The Moderation AI Agent uses powerful AI models to evaluate content in real time, helping moderation teams screen posts faster and more consistently.

 

Why AI Moderation?

 

Moderation AI Agent helps enforce your community’s code of conduct by:

  • Automatically detecting potentially inappropriate or harmful content.
  • Flagging submissions for manual review or blocking them automatically based on risk level.
  • Complementing existing tools such as Keyword Blocker and Spam Prevention.

Moderation AI Agent can reduce the manual workload of your moderators, improve consistency, and accelerate content verification, while continuing to provide a safe and high-quality experience for all community members.

 

How Does AI Moderation Work?

 

Moderation AI Agent classifies posts and replies into Approved, Pending, or Trash, applying informative reasoning and descriptive moderator tags for human intervention, filtering, and review.

It provides a fully automated moderation by immediately processing all User-Generated Content (UGC) through OpenAI's Moderation API, ensuring initial compliance. Post-clearance, the AI Moderation further validates content against detailed checks for adherence to community guidelines, absence of NSFW content, PII protection, and spam detection.

 

Community Code of Conduct

 

The code of conduct provides guidelines for the Moderation AI Agent and sets clear guardrails for moderation. This ensures that AI-powered moderation aligns with the unique needs of your community.

You can reuse your existing public-facing community rules, code of conduct, or equivalent guidelines. In addition, you may include internal training documents that community managers use during onboarding. The Moderation AI Agent evaluates content against your code of conduct to make the initial decision about what is acceptable within your community.

Note: The Community Code of Conduct can be up to 5,000 characters in length.
 

Moderation Status

 

Moderation AI Agent scores content on a scale of 0.0 - 1.0 and currently has three possible outcomes:

Status

Confidence Score

Description

Approved

0.0 - 0.4

Content is considered safe and appropriate. Content is approved, published, and visible in the community

Pending

0.4 - 0.7

Requires manual review; borderline or uncertain content. Content is held in the Pending status, is not published, and is not visible in the community.

Trash

0.7 - 1.0

Content violates community or Gainsight moderation guidelines. Content that is Trashed or Trashed and Reported is not published and is not visible in the community.

 

Moderator Tags

 

Moderation AI Agent adds tags to each topic or reply that it reviews, to share insights on sorting, filters, and other analytics based on all content moderated.

 

Category

Tags

Positive Case

  • Meets code of conduct
  • SFW
  • No PII
  • No Spam
  • Approved

Pending Case

  • Pending code of conduct
  • Pending NSFW
  • Pending PII
  • Pending Spam
  • Pending

Negative Case

  • Does not meet the code of conduct
  • NSFW
  • Contains PII
  • Contains Spam
  • Trash

Other Cases

  • Flagged by OpenAI
  • AI Moderator
  • To be reviewed by the community manager

 

Configure AI Moderation

 

Moderators can configure Moderation AI Agent in addition to Keyword Blockers and Moderator Approval to ensure that the AI evaluates the content before it is published in the community.

To configure AI Moderation:

  1. Log in to Control.
  2. Navigate to AI > Moderation AI Agent
  3. Turn on the Use AI Moderation toggle.
    Note: When AI Moderation is enabled, an AI moderator user is created. Gainsight recommends not deleting this user.

     

  4. In the Code of Conduct, enter your community guidelines.
  5. (Optional) In addition to Administrator, Community Manager, Moderator, and Superuser, you can add additional roles whose content can be excluded from AI Moderation

  6. (Optional) Enable User profile moderation to review New member registrations. Registrations that don't pass are held in Pending for moderator review. Administrators, Community Managers, Moderators, and Superusers are exempt.

  7. Click Save changes.

Content Moderation Widget

 

Once Moderation AI Agent is configured, any Topics or Replies that do not meet your community’s guidelines are automatically tagged with Moderator Tags. These posts are then moved to Trash and Reported.

 

You can review this content in the Content Moderation widget on the Control Home page.

For more information on how to add this widget, refer to the Overview of Control Home article.

    44 replies

    revote
    VIP ⭐️⭐️⭐️⭐️⭐️
    October 7, 2025

    Okay, so how this actually works?

    (Sorry, I was not able to delete those tables here)

    How Does AI Moderation Work?

     

    Moderation AI Agent classifies posts and replies into Approved, Pending, or Trash, applying informative reasoning and descriptive moderator tags for human intervention, filtering, and review.

         

     

    Moderator Tags

     

    Moderation AI Agent adds tags to each topic or reply that it reviews, to share insights on sorting, filters, and other analytics based on all content moderated.

       

     

     

    Suvi Lehtovaara
    Helper ⭐️⭐️
    October 8, 2025

    Hi ​@Graeme Rycyk , 

    I am testing this tool in our sandbox - seems to work well even in Finnish language 💪🏻

    However I have found some cases where synonyms like nasty, mean, vicious, unkind were not always recognised by the tool. How can we teach the AI?

    https://yhteiso.elisa.fi/
    Suvi Lehtovaara
    Helper ⭐️⭐️
    October 8, 2025

    And another thing came up:

     


    The upper comment/reply contains a phone number > it was trashed.
    The lower one contained an email address > it was automatically moderated.

    Why the logic is different? Is it because email is more easily recognized as an email?

    https://yhteiso.elisa.fi/
    Graeme Rycyk
    Gainsight Employee ⭐️
    Gainsight Director of Product
    October 8, 2025

    Hey ​@Suvi Lehtovaara,

    An easy way to solve this would be to add words you want to ensure are always caught into your Code of Conduct.

    Do let me know if you have any other questions or if this doesn’t work.

    Kind regards,

    Graeme

    Director of Product | AI & Search
    mdfahd
    mdfahdAuthor
    Gainsight Employee ⭐️⭐️
    October 8, 2025

    Hi ​@Suvi Lehtovaara,
    Mohammed Fahd here, I’m the POC from documentation for this article. We have recently received a request from you to access the Overview of Control Home article. Apologies for the inconvenience as the link was wrongly tagged. 
    We have now updated the article link - Overview of Control Home. Please let me know if you have any trouble accessing it. Thank you

    Mohammed Fahd - Sr. Technical Writer
    revote
    VIP ⭐️⭐️⭐️⭐️⭐️
    October 9, 2025

     

     

    Thanks, this is helpful post.

    So there is no Moderation tags in replies, so we dont know what and why actions are made.

    Suvi Lehtovaara
    Helper ⭐️⭐️
    October 9, 2025

     

     

    Thanks, this is helpful post.

    So there is no Moderation tags in replies, so we dont know what and why actions are made.

    Yes, that’s true ​@revote. And if there’s no context in the post, the reason may stay unclear.

    https://yhteiso.elisa.fi/
    Graeme Rycyk
    Gainsight Employee ⭐️
    Gainsight Director of Product
    October 9, 2025

     

     

    Thanks, this is helpful post.

    So there is no Moderation tags in replies, so we dont know what and why actions are made.

    At this moment yes, this is correct, but this is an improvement we intend to make in the near future.

    Kind regards,

    Graeme

    Director of Product | AI & Search
    revote
    VIP ⭐️⭐️⭐️⭐️⭐️
    October 9, 2025

    Yes, that’s true ​@revote. And if there’s no context in the post, the reason may stay unclear.

    Also, I didn’t know that AI deletes content automatically. I thought it just hides content that goes against the code of conduct.

    Why is this a problem?

    If you think about the DSA, the Digital Services Act: if we limit user communication by deleting content or even just part of it, the DSA states that the user has the right to make a note about the decision (the decision to delete content). And we are obligated to handle that note.

    In theory, the decision to delete might be wrong. That means we would have to restore the content.

    But since there is no version history in replies, restoring the content is not possible.

    We need to consider whether we should start using Moderation AI at this point.

    Graeme Rycyk
    Gainsight Employee ⭐️
    Gainsight Director of Product
    October 9, 2025

    Yes, that’s true ​@revote. And if there’s no context in the post, the reason may stay unclear.

    Also, I didn’t know that AI deletes content automatically. I thought it just hides content that goes against the code of conduct.

    Why is this a problem?

    If you think about the DSA, the Digital Services Act: if we limit user communication by deleting content or even just part of it, the DSA states that the user has the right to make a note about the decision (the decision to delete content). And we are obligated to handle that note.

    In theory, the decision to delete might be wrong. That means we would have to restore the content.

    But since there is no version history in replies, restoring the content is not possible.

    We need to consider whether we should start using Moderation AI at this point.

    Hey again ​@revote,

    So AI Moderation doesn’t delete any content, it only Trashes and adds a Report reason for further context for content that was trashed. This is how content is removed from public viewing in our communities. If mistakes are made by the AI Moderator, say by mistakenly Trashing content, this can always be reversed by the human-moderator. Our “Notify Author” flow can be applied by your community team to adhere to the DSA even if using AI Moderator. Custom views can be set up to enable your community team to easily create a workflow to track the moderation action of the AI Moderator.

    I would be happy to jump on a call to discuss things and take more questions, if you would prefer.

    Kind regards,

    Graeme

    Director of Product | AI & Search