BrandBastion Blog

How to Set Up Twitch Moderation and Keep Chat Safe

Written by BrandBastion | 8/23/26, 10:55 PM

Your chat used to fit on one screen. Now it scrolls faster than you can read it while you are also running the stream, and somewhere in that scroll a troll is testing what they can get away with in front of your community.

That is the moment Twitch moderation stops being a setting you flipped on once and becomes a job. Left unmanaged, the cost is not abstract: harassment sits visible under your community's nose, gets screenshotted, and does the reputational damage before you ever see it.

By the end of this guide you will have a layered moderation setup on Twitch, from AutoMod through a real mod team, and a clear read on where that setup stops scaling.

Most streamers treat moderation as a checklist they configure once and forget. The mistake is assuming more tools or more mods alone will hold the line at volume. They will not: the tools are layers that only work together, and manual, volunteer moderation has a ceiling you will eventually hit.

In this guide

  1. Before you start
  2. Turn on your baseline chat protections first
  3. Set your AutoMod level and understand what it actually does
  4. Add your own blocked and permitted terms
  5. Add moderators and set what mods and VIPs can do
  6. Equip your mods with Mod View and the core actions
  7. Lock chat down fast during a raid or harassment spike
  8. Extend native tools with third-party chat bots
  9. Build and run a mod team that stays consistent
  10. When manual moderation stops keeping up
  11. Do this today

Before you start

You need three things in place before you configure anything:

  • Broadcaster access to your channel's Creator Dashboard, where every setting below lives.
  • A verified account (email, and phone if you can) so the verification-gated chat modes actually work when you need them.
  • A rough decision about who, if anyone, will help you moderate.

A note if you are a team of one: some steps below assume headcount you may not have, like a full mod team or 24/7 coverage. Where that happens, the scaled-down version is called out. You can run a solid setup solo. You just lean harder on automation and chat modes to cover the gap.

Turn on your baseline chat protections first

Before you touch anything advanced, set the floor. This is the layer that stops the most common low-effort spam and drive-by trolling before it reaches chat, and it takes a few minutes.

Set these three:

  • Account verification for chatters, so throwaway accounts cannot post freely.
  • Follower-only mode, which restricts posting to accounts that have followed your channel for a minimum time you set (a useful brake against fresh accounts spun up to spam).
  • AutoMod, switched on at a starting level (the next step covers how to pick one).

Treat this as layer one of a system, not a fix on its own. It handles the obvious noise so the tools above it can focus on the harder calls.

Set your AutoMod level and understand what it actually does

AutoMod is Twitch's built-in language filter. It uses automated detection to catch potentially risky messages and hold them back, so a human can approve or deny each one before it ever appears in chat. That last part is the whole point, and it is the part most guides skip.

You control how aggressive it is with a single slider. Per Twitch's AutoMod filtering levels, there are five levels of filtering: 0 (no filtering) through 4 (strongest filtering), applied across categories including aggression, sexuality_sex_or_gender, misogyny, bullying, swearing, race_ethnicity_or_religion, and sex_based_terms. At every level, AutoMod withholds flagged messages for a moderator to approve or deny before they appear in chat.

The human-in-the-loop design is not new. In AutoMod's original 2016 design, published December 12, 2016, streamers could configure AutoMod by selecting one of four levels across identity, sexual language, aggressive speech, and profanity, and when AutoMod flags a message, moderators review its content before it is sent to chat. The category list has grown since. The approve-or-deny model has not.

Here is the correction worth internalizing: AutoMod does not time out, ban, or mute anyone. It only withholds messages pending review. It is a shield, not a punishment. The timing out and banning is still a human decision, which means AutoMod does not reduce the need for moderators. It changes what they spend their attention on, moving them off the obvious slurs and onto the judgment calls.

Start in the middle and adjust from what you actually see. A mid-level setting catches overt hate and slurs while letting normal conversation through, which is the right default for most channels finding their footing.

Then watch for the two failure modes, because you will hit one of them.

Set it too low and borderline harassment slips through. The tell is that you or a mod keep manually deleting messages AutoMod should have caught, or viewers start reporting things you never saw. Nudge the level up.

Set it too high and the approval queue backs up. The tell is legitimate chat stalling, held messages piling up faster than anyone can clear them, and regulars asking why their messages are not showing. Every held message needs a human to release it, so an over-tuned AutoMod does not remove work. It creates a bottleneck. Nudge it down, or use permitted terms (next) to stop it flagging the specific words that are clogging the queue.

Add your own blocked and permitted terms

AutoMod is general-purpose. Your community is specific. Two lists close the gap.

Blocked terms let you ban words and phrases AutoMod would miss: a slur being aimed at your specific community, a targeted harassment phrase tied to a recent incident, or spam URLs making the rounds. Permitted terms do the reverse, letting through words AutoMod keeps holding that are harmless in your context.

As an illustration, say your community has an inside-joke word that reads as borderline out of context, and AutoMod flags it every stream. Rather than lower your whole filter level, you add that one word to permitted terms. The meme flows, and the filter stays strict on everything else.

Treat both lists as ongoing maintenance, not a one-time setup. New slang, new spam, and new harassment patterns show up constantly. These lists are the manual layer that patches the automated one, which is exactly why they need you to keep tending them.

Add moderators and set what mods and VIPs can do

When you have someone you trust, promote them to moderator. A moderator can delete messages, time users out, ban, and manage your chat modes while you focus on streaming.

Do not confuse a moderator with a VIP. A VIP is a trusted regular who bypasses chat restrictions like slow mode and follower-only mode, but has no moderation powers. VIP is a recognition-and-access role; moderator is a responsibility. Give VIP to loyal community members you want to reward. Give moderator only to people whose judgment you would trust to act in your name when you are mid-stream and not watching chat.

Pick mods who are present when you stream, level-headed under pressure, and aligned with how you want your community to feel. Skills you can teach. Temperament you cannot.

If you are a team of one, you may have zero mods at first, and that is fine. The scaled-down move is to lean harder on AutoMod and your chat modes until a regular earns the role by showing up consistently and reading the room the way you would. Do not hand the keys to someone just to fill the slot.

Equip your mods with Mod View and the core actions

Mod View is the dedicated interface your moderators work from while you stream: a moderation-focused layout that puts chat, user info, and quick actions in one place, separate from the normal viewer experience. It is where the actual work happens.

The core actions are three:

  • Delete removes a single message. Reach for it on a one-off that does not warrant touching the person.
  • Timeout mutes a user for a set duration. Use it as a warning or a cooldown, a way to interrupt someone who is heating up without permanently removing them.
  • Ban removes a user for good. Save it for repeat offenders and clearly over-the-line behavior.

The harder problem is not the actions. It is consistency across the people taking them. Two mods enforcing the same rule differently confuses your community and undermines both of them. This is why the audit trail matters: being able to see what actions your mods have taken lets you spot drift, coach it, and keep enforcement even. That consistency challenge only grows as you add people, which is what the mod-team step is about.

Lock chat down fast during a raid or harassment spike

Everything above is your standing setup. This is what you do when chat is being flooded right now, whether by a raid (a wave of accounts hitting your channel at once, sometimes friendly, sometimes coordinated to harass) or a hate raid built specifically to overwhelm your moderation.

You have two layers to reach for: manual chat modes and Shield Mode.

The chat modes let you dial up restrictions instantly:

  • Slow mode limits how often each user can post, throttling a fast-scrolling flood.
  • Follower-only mode restricts posting to established followers, cutting out accounts created minutes ago.
  • Subscriber-only mode limits chat to paying subscribers, a harder gate for when a spike is severe.
  • Emote-only mode allows only emotes, which shuts down text-based slurs and spam entirely while keeping chat alive.
  • Non-mod chat delay adds a short buffer before non-moderator messages appear, giving your mods a window to catch and remove something before viewers see it.

Stacking those by hand mid-incident is slow, which is why Twitch built a single control for it. Twitch's one-click Shield Mode, published November 30, 2022, is a one-button toggle that can limit chat to followers or subscribers, require verification and implement stricter AutoMod levels, or automatically ban everyone who recently used a given phrase, then immediately revert back to looser policies once the crisis is over. You configure the bundle in advance and fire it in one action when you need it.

When a raid or hate raid is hitting right now, work in this order:

  1. Hit Shield Mode. If you have not set it up yet, switch on emote-only or subscriber-only mode immediately. Stop the bleeding before you do anything else.
  2. Tell your mods to ban and purge the incoming accounts rather than argue with them. During a coordinated attack, engagement is what the attackers came for.
  3. If the raid is riding a specific phrase or slur, use the ban-by-phrase option to clear everyone who used it in one sweep.
  4. Once the wave passes, revert your settings so real chat can breathe again. Shield Mode reverses the same way it turned on.
  5. Report and document. Screenshot the worst of it before you purge, then report the originating channel or accounts.

That last step matters more than it looks. Coordinated raids are not only a chat-settings problem. They are a policy violation. Under Twitch's Hateful Conduct Policy, effective January 22, 2021, the policy evaluates the content of statements or actions over perceived intent, and explicitly prohibits inciting malicious raids of another person's social media profiles off Twitch. Reporting is part of the response, not an afterthought, because the fix for a coordinated attack often sits above your channel's own settings.

Extend native tools with third-party chat bots

Third-party chat bots add a layer on top of what Twitch gives you natively: more granular link and spam filtering, custom automated commands, timed messages, and rules you can tune beyond AutoMod's slider. Some of what they do now overlaps with native features that have caught up over the years, so audit honestly before adding one, and do not run a bot for a job AutoMod already handles.

Treat a bot as one more layer, not a replacement for AutoMod or your human mods. It automates repetitive enforcement so your mods spend less time on obvious spam and more on the calls that actually need a person.

Build and run a mod team that stays consistent

A mod team is an ongoing management practice, not a one-time hire. Get it right and enforcement stays even whether you are watching or not. Get it wrong and you have several people improvising different rules in public.

Twitch runs on this labor more than most people realize. A peer-reviewed study of Twitch moderators (Seering & Kairam, "Who Moderates on Twitch and What Do They Do?", Proc. ACM Hum.-Comput. Interact., Jan. 2023) found that, per Twitch's H1 2021 Transparency Report, volunteer moderators and user-created bots covered 87.8% of live minutes watched, removing more than 37 million messages manually in that period. In the paper's survey of 1,053 active moderators, 68% work in teams of three or more.

A few principles for running one:

Size to volume, not to a round number. Match your mod count to your concurrent chatters and your streaming hours. The signal you need more is simple: messages that should have been actioned sit in chat because no one had the capacity to catch them.

Write the rules down. A short, shared document of what earns a delete, a timeout, and a ban is what makes several different people enforce like one. Without it, consistency is impossible no matter how good each mod is individually.

Define escalation. Everyone should know what to handle themselves, what to bring to you, and what triggers a lockdown. Ambiguity in the moment is how incidents get worse.

Watch for burnout. Moderation is emotional labor, and your mods are absorbing the abuse so your viewers do not have to. Rotate coverage, thank them, and do not expect volunteers to provide round-the-clock protection.

And be honest about scale here, because this is where the whole approach shows its ceiling. Adding people does not add capacity in a straight line: every new mod is another person to onboard, align, and keep consistent, and volume can still outrun the entire team during a spike.

If you are a team of one, you are not failing because your team is you plus one or two trusted regulars leaning on automation. That is the legitimate scaled-down version. Write your rules down anyway, even if the only person reading them today is you, because the document is what makes handing off possible later.

When manual moderation stops keeping up

AutoMod plus a volunteer team is a strong setup, and it still has a boundary. It is worth naming plainly so you recognize the edge when you reach it.

Manual moderation breaks down in predictable places: message velocity outruns human review during spikes, coverage gaps open after hours and across languages your mods do not speak, fatigue sets in, and enforcement drifts between people. None of that is a discipline problem you can train away. It is structural. Past a certain volume, human review simply cannot keep pace, and unmanaged high-volume conversation becomes a real, measurable cost: wasted moderator effort, harassment seen before it is caught, and eroded community trust.

That same logic, that conversation at scale is a cost rather than a nuisance, is what BrandBastion addresses on brand-owned social surfaces: the comments, DMs, mentions, and reviews across Instagram, Facebook, TikTok, X, LinkedIn, and YouTube. That is a different problem from moderating live Twitch chat, and the approach pairs automation with human oversight rather than handing judgment to a machine. For a sense of what managing high-volume conversation well looks like as an outcome, JD Sports' engagement results are one example.

Do this today

Do not wait until an incident forces the setup. Before your next stream, turn AutoMod on at a sensible middle level and switch on follower-only mode. That puts the baseline layer live in about five minutes. Then decide one thing: who your first moderator will be, or, if no one has earned it yet, which chat modes you will keep on until someone does.

You will know your Twitch moderation is working when the approval queue stays manageable and harassment gets caught before a viewer can screenshot it. Everything else in this guide is you adding layers toward that state, and knowing which layer to reach for when the last one stops keeping up.