Two professionals discussing AI strategy in a collaborative setting.
Go back

Artificial Intelligence & Data, Innovation & Emerging Technology, Privacy and security

When AI escapes the sandbox: From public conversations to rogue agents

Indexing incidents signal the need for a more mature conversation about responsible AI at scale, including data privacy and Trust and Safety. Rob Lewington, Senior Vice President, Head of Trust and Safety at TP - 9/9/2026

Recently, public AI safety incidents, from supposedly private AI conversations that become searchable online to unauthorized data exposures and database hacks, forced the industry to resolve questions it has long deferred. This is a familiar pattern for anyone who has worked through earlier waves of digital innovation.

 

When new technology comes along, the first principles are often the development of core tech and getting it onto the market as quickly as possible. That process of first optimizing the product then ensuring product-market fit before focusing explicitly on safety is often inherent to the nature of our society. Consider the counter-factual: developing an explicitly safe product that is not compelling to use would likely not be a successful business strategy.

 

 

However, what we learned with the advent of user-generated content (UGC) platforms like social media and gaming is that product safety is a core part of what customers, governments, and society-at-large expect from technology companies.

 


Strong Trust and Safety practices protect user data and boost confidence in technology

 

The patterns we are experiencing today with AI follows a clear trajectory:

 

  • New technologies reach the market and the immediate priority becomes access and adoption 
  • Growth accelerates as businesses move quickly to capture the opportunity 
  • Safety concerns become more visible once these systems are already reaching large audiences 
  • Public confidence can weaken when the safeguards behind the technology are unclear 

 

It is no coincidence that public trust in AI is already under pressure.

 

According to recent data from the Pew Research Center¹, 50% of Americans report feeling more concerned than excited about the increased use of AI in daily life, with only 10% feeling more excited than concerned. Together, these findings reinforce a broader point: trust is hard to build and easy to lose.   


Learning from the past to protect the future

 

A lack of trust in social media has developed over time. As platforms scaled, safety infrastructure and content moderation often struggled to keep pace with the volume of UGC. Harmful material could remain visible for too long, and users began to question whether the platforms they relied on were designed to protect them. 

 

AI is a complex, generational technology evolving at an unprecedented pace, so challenges are inevitable. The opportunity now is to apply lessons from the past early by building safety into the design and operation of AI systems as they scale. This is where Trust and Safety becomes essential. It gives companies the human expertise and operating framework needed to identify risks, protecting both the user experience and confidence in the technology, grounded in a profession that has come to maturity during the social media era. 

 

This is in the interests of all parties because the social media parallel now has tangible implications for user behavior. In its 2024 Trust Barometer³, Edelman, a global communications firm, found that trust in social media platforms ranked at or near the bottom among business sectors, dipping as low as 24%. That distrust also correlates with 55% of Americans reporting that they post less frequently than they did five years ago. 

 

The positive aspect from the conversations I am having is that AI companies are starting to recognize this, and I would urge companies to keep thinking about safety as a business imperative. 
 

Why Trust and Safety must be built into AI systems

 

One recurring mistake I have seen throughout the era of UGC and one I hope AI companies will avoid is the effort to continually try to “design out” the need for content moderation or other Trust and Safety work. This can take the form of crowdsourcing Trust and Safety entirely through user moderation, creating blocking and muting tools in the hope that they will be sufficient to maintain a positive ecosystem, or relying too heavily on automation. 

 

These are undeniably important parts of the toolkit. They do not replace the need for a robust Trust and Safety foundation, and the investment in that foundation must be recognized as part of the cost of doing business. This will save product teams significant time and money down the line, because they will not have to make fundamental changes to the core infrastructure later. 

 

How AI safety works in practice

 

Across the systems I have worked on, strong Trust and Safety frameworks rest on a few practical disciplines:

 

  • Before teams commit to product architecture, they assess the likely harms that could follow 
  • Human specialists remain accountable for quality as AI takes on greater review volume
  • A clear record of where data came from and how it changed makes risk reviews easier 
  • Continuous monitoring matters because models and threats can drift as systems scale 

Making the business case for proactive safety

 

During the last wave of Trust and Safety work, teams often designed features in what I would call a “minimum viable product” way. When safety considerations are not fully scoped early, predicted worst-case scenarios can force engineering teams to spend cycles unpicking an avoidable problem. AI coding and development tools may reduce the time lost to this kind of scope narrowing, but it still takes development resources away from other priorities. 

 

To be clear: considering Trust and Safety upfront does not necessarily mean building all safety protocols necessary for every conceivable eventuality or scenario, as you must also be realistic and stay agile. But it does mean doing proper, thorough risk assessments and making informed design decisions with these potential consequences and trade-offs in mind. 

 

The same discipline applies to the data behind an AI system. In a Forbes Technology Council article⁴, Akash Pugalia, Chief Digital Officer at TP, writes that companies have not always checked closely enough “where the data came from and whether it can be trusted.” Keeping a clear record of how data is collected and changed makes it easier to assess risks and respond when an issue arises. 


Translating frameworks into tangible value

 

What separates a meaningful framework from a cosmetic one ultimately comes down to real-world outcomes and the ability to quantify them. As someone who was a TP client before joining the company, I believe the full depth and breadth of what we do in Trust and Safety is not always visible from the outside.

 

At TP, we have more than 40,000 full-time professionals dedicated to Trust and Safety work at any time. Together, they make nearly 3 billion Trust and Safety decisions annually. The scale of that is really quite extraordinary.

 

A piece of research that puts this in sharp focus used the public ledger of the DSA, the European Union’s Digital Services Act⁵, a transparency database to measure how quickly platforms remove illegal content. Their simulations found that removing content within two hours reduces total societal exposure to it by over 70%, but that longer delays were measured across major platforms, failing to contain its spread.

 

This is an area where effective AI-implementation can drive improved safety outcomes, since vast amounts of content can be reviewed with significantly reduced response times. This represents the continued shift of content moderation from large teams of “human-in-the-loop” (HITL) reviewers to specialized “human-on-the-loop” (HOTL) experts who ensure this speed of response does not come at the cost of reduced quality.


What makes a Trust and Safety framework effective at scale

 

AI solutions underpinned by deep, specialized human expertise push the needle on problems like reducing exposure to harmful content, and when we see those kinds of quantified results, we will know we are doing something right. That is exactly what TP is mobilizing its scale to help solve.

 

Trust and Safety is part of TP.ai Dataservices, bringing data expertise and human oversight to the work of building AI systems that can scale responsibly. In April 2026, TP.ai Dataservices was named Data Analytics Platform of the Year by the Data Breakthrough Awards, recognizing TP's work at the intersection of AI and human expertise to produce outcomes clients can measure. TP also continues to build its responsible AI governance foundation through BSI certification for AI management, reflecting the operational standards and trust that global transformation partners are expected to maintain.

 

Talk to us to learn how TP can help you build and scale a robust Trust and Safety framework for your AI operations.




Other impactful stories

Insightful Articles
  • Insightful Articles