Roblox Shares New Open-Source Safety Tools With ROOST

Roblox Open-Sources New Safety Models for PII, Child Protection and Voice Chat

Roblox is expanding access to its online safety technology by contributing three open-source models to the Robust Open Online Safety Tools, or ROOST, Model Community. The release includes an updated PII Classifier, Roblox Sentinel and the company’s latest voice safety classifier, along with a new evaluation dataset for testing personal information detection systems.

Versions of these technologies are already used on Roblox to identify attempts to request or share personal information, detect early indicators of potential child endangerment and moderate inappropriate speech in voice chat.

Roblox joined ROOST as a founding member in early 2025. By making the models available to the wider community, the company aims to give other platforms a foundation they can adapt and improve for their own moderation systems.

The August 19, 2026 announcement also provides new performance data for the technologies, including a substantial improvement to Roblox’s PII Classifier and expanded multilingual support for voice moderation.Roblox Shares New Open-Source Safety Tools With ROOST

Roblox Updates Its PII Classifier With Support for 189 Languages

One of the largest updates concerns the Roblox PII Classifier, a system designed to detect attempts to share or request personally identifiable information.

The classifier is also intended to identify attempts to move users to other platforms, including cases where people try to avoid moderation through misspellings, coded language or indirect references.

Instead of relying only on exact words or individual messages, version 2.0 can examine surrounding conversational context. This is important because a single message may appear harmless when viewed independently but take on a different meaning when considered alongside earlier messages.

Roblox has also expanded the model’s language coverage significantly. Version 2.0 supports 189 languages, compared with 17 previously.

The expansion was supported by synthetic training examples generated with large language models, including more complicated attempts to bypass moderation and less common multilingual cases. Roblox also paired LLMs with context retrieval from curated labeled data.

According to Roblox, these changes improved the classifier’s F1 score from 63.41 in version 1.0 to 90.52 in version 2.0.

Safety Technology Primary Purpose Key 2026 Update
PII Classifier v2 Detect PII sharing, solicitation and attempts to move users off platform 189 languages, conversational context, F1 score of 90.52
Roblox Sentinel v2 Identify early signals of potential child endangerment More score-combination methods and faster configuration testing
Voice Safety Classifier v3 Detect inappropriate speech in voice chat 30 languages, eight violation categories and built-in language detection
PII Safety for Chat Benchmark Evaluate PII classifiers Synthetic multiuser conversations designed around realistic evasion techniques

New Benchmark Tests How PII Filters Handle Evasion Attempts

Alongside the updated classifier, Roblox is releasing a new evaluation dataset called the Roblox PII Classifier Benchmark.

The dataset contains synthetic English-language multiuser conversations designed to reproduce techniques that can be used to avoid conventional PII filters. The simulated conversations include users requesting or sharing personal information and attempting to direct other users away from the platform.

The benchmark covers several types of evasion, including:

  • Phonetic variations
  • Character and visual substitutions
  • Coded references
  • Information divided across multiple messages
  • Attempts to direct users to other platforms

Roblox also included hard negative examples. These are benign conversations that resemble potentially problematic interactions but do not actually contain PII solicitation or sharing.

This distinction is important for moderation systems because detecting suspicious-looking words alone can produce unnecessary false positives. Context can help determine whether a conversation represents an actual attempt to exchange personal information.

Roblox says many existing datasets focus on named-entity extraction, while real online conversations can involve attempts to solicit information before an identifiable email address, username or other PII string ever appears.

The new benchmark is intended to allow developers and online platforms to evaluate their own classifiers against these more conversational scenarios.

Roblox Sentinel Targets Early Signs of Child Endangerment

Roblox is also releasing version 2 of Sentinel, its open-source system for identifying early signals of potential child endangerment.

Sentinel uses contrastive learning and is trained on patterns found in both benign conversations and conversations that eventually become harmful. It compares new interactions with those patterns to identify subtle signals that may warrant additional review.

The objective is to detect concerning behavior before a conversation becomes explicit enough to trigger more conventional moderation systems.

Sentinel helps human reviewers prioritize conversations that may require attention, allowing them to investigate potentially harmful behavior and take appropriate action.

Roblox reports that, during the 12 months ending August 7, 2026, nearly 70% of the cases it detected were identified through Sentinel’s early detection capabilities.

That figure does not mean Sentinel identified 70% of all harmful activity on Roblox. It specifically refers to the cases detected by Roblox during the period described by the company.

Sentinel Version 2 Improves Accuracy and Testing Speed

The new version of Sentinel introduces several technical improvements intended to make the open-source system easier to configure for different applications.

Version 2 expands the number of available score-combination functions from two to six. Different functions can be selected depending on the type of data being analyzed.

Roblox has also added tools for evaluating which function is most appropriate for a particular use case, including explanations of why a source received a particular score.

Another change significantly reduces the computational work required when testing configurations. The new version eliminates the need to recompute evaluation data for every setting.

In one example published by Roblox, Sentinel version 2 completed a sweep across 324 configurations in less than three minutes. The same process would previously have taken nearly an hour.

Roblox also reports that ROC-AUC, a metric used to measure ranking performance, increased from 0.894 under the default settings to 0.996 using the best configuration identified during testing.

These improvements are intended to make it more practical for other organizations to experiment with Sentinel and adapt it to their own datasets and safety requirements.

Voice Safety Classifier Expands to 30 Languages

Roblox’s third major open-source contribution is version 3 of its voice safety classifier.

The system is designed to detect inappropriate speech during voice interactions in real time. Roblox first open-sourced the classifier in 2024, and the company says it has since been downloaded more than 72,000 times.

Version 3 supports a total of 30 languages and eight violation categories. It also includes built-in language detection.

Across the 30 supported languages, Roblox reports 61% recall at a 1% false-positive rate. The company says improvements in precision and recall were supported by a larger amount of labeled training data, combining machine labeling for scale with human labeling for quality.

The underlying model has also grown considerably. Its size increased from 94.6 million parameters to 320 million parameters.

Despite the larger model, Roblox says it used model distillation to maintain the speed required for real-time moderation.

How Roblox Uses the Voice Classifier During Gameplay

The voice safety classifier is connected directly to Roblox’s moderation process.

When a player violates a rule, Roblox can display a notification explaining that a policy has been broken and identifying the relevant policy. Repeated violations can result in temporary suspension of voice chat for up to five minutes.

More serious violations or reports from other players can result in additional consequences.

This approach allows Roblox to identify some policy violations as voice interactions happen and respond within the platform rather than relying entirely on reports submitted after an interaction.

Why Roblox Is Sharing Its Safety Technology Through ROOST

Roblox became a founding member of ROOST in early 2025 alongside organizations including Google and OpenAI. The initiative focuses on making online safety technology more accessible through open-source tools and collaboration.

Roblox’s contribution gives other organizations access to technologies developed around moderation challenges encountered on a large online platform.

The three systems address different parts of online safety. The PII Classifier focuses on personal information and attempts to move conversations off platform. Sentinel is intended to detect early warning signals associated with potential child endangerment. The voice classifier focuses on inappropriate speech during real-time voice interactions.

By open-sourcing the models, Roblox aims to give other platforms a starting point for training and tuning moderation technology for their own use cases.

Roblox also expects the exchange to work in both directions, with feedback and developments from other ROOST participants potentially contributing to future improvements.


Roblox Expands Its Open-Source Approach to Online Safety

The latest releases broaden Roblox’s open-source safety program across text, behavioral risk detection and voice communication.

The PII Classifier now uses conversational context and supports 189 languages, Sentinel version 2 provides more flexible and significantly faster configuration testing, and the voice safety classifier supports 30 languages while remaining suitable for real-time detection.

The accompanying PII benchmark adds another element by giving developers a dataset specifically designed to test classifiers against conversational attempts to evade moderation.

For Roblox users, these technologies operate behind the broader experience of playing, communicating and creating on the platform. Players who use Roblox and want to add Robux to their accounts can also find Robux Global vouchers through Baxity Store. Available denominations and current stock can vary, so users should check the product details and redemption requirements before purchasing.

Leave a Reply

Your email address will not be published. Required fields are marked *

Related

The Baxity.com website in any way does not promote gambling, betting, or any other services that have legal, age or other restrictions and require licenses for the companies providing these services and does not encourage users and any persons to use any of these services. Any materials available on the website are fact-finding articles for users of electronic payment systems that are regulated by the relevant supervisory authorities of the Republic of Estonia, the European Union and Saint Vincent and the Grenadines. If the legislation of your country prohibits the use of this kind of content or services, or you have not reached the age of majority, then refrain from using our website.