Boosting Trust and Safety in Generative AI

By Tanya Sattineni

In March 2023, an AI-generated image of Pope Francis wearing a Balenciaga jacket spread around social media, garnering over 20 million views and prompting experts to call it one of the first mass cases of AI-generated misinformation. This incident exemplifies a looming challenge: Generative AI makes misinformation faster, cheaper, and increasingly convincing. 

Failures across major AI firms show that consequences of misinformation are already real. Tools from OpenAI have generated fabricated legal citations, and actors have used them in influence operations on social media platforms, while hackers have exploited Anthropic systems to run phishing scams. From the surge of AI-generated election misinformation and deepfakes in 2023 to high-profile failures and controversies involving Google’s Gemini AI in 2024, public trust in AI companies dropped from 61 to 53 percent globally. Technology companies must preemptively mitigate AI misuse to stay competitive. 

The stakes extend beyond online deception. AI-generated misinformation can erode trust in institutions, undermine democratic processes, and intensify societal polarization. Synthetic images, voices, and narratives blur the line between authentic and artificial information. Media consumers face an increasingly difficult task of judging credibility and accuracy. At the geopolitical level, AI tools can amplify information warfare, allowing adversaries to rapidly deploy influence campaigns that manipulate public perception and shape policy debates. 

Despite growing discourse on the malicious use of AI, private industry governance still falls short. Safeguards are typically reactive, not proactive. Following multiple incidents in which lawyers submitted hallucinated legal citations generated by ChatGPT, OpenAI introduced stricter policies limiting legal advice and emphasizing safeguards. 

Meanwhile, OpenAI has entered into partnerships with government entities like the U.S. Department of Defense. Prematurely implementing software that tends to fabricate information raises serious concerns for potential state-led surveillance and military applications. 

Some progress exists: OpenAI, Google, and Anthropic have expanded red-teaming to uncover vulnerabilities, while Google’s SynthID and Anthropic’s transparency reports provide some traceability. However, these measures are uneven. For example, watermarking for traceability is not universally adopted across companies. As a result, AI systems remain vulnerable, limiting the ability to build durable public trust. 

Comprehensive AI governance requires collaboration not only among AI developers but also between companies and governments. Intelligence sharing can help identify emerging misinformation tactics early. At the same time, both AI developers and digital platforms must implement safety tools: Developers can embed toxicity filters and content safeguards within models, while platforms like YouTube and Facebook can deploy deepfake detection and moderation systems to limit the spread of misleading content. 

However, these interventions face serious practical limitations. Sharing threat intelligence across competitors is often constrained by privacy concerns, antitrust laws, and proprietary interests. Synthetic content detection can often produce false positives or be evaded by sophisticated adversaries. As a result, while these safeguards are conceptually valuable, their real-world effectiveness is likely to remain uneven. 

However, end-to-end content traceability offers a more proactive approach. By embedding watermarks or cryptographic markers at the point of creation and tracking how content circulates, companies can more reliably identify AI-generated material and intervene when misuse occurs. Private companies can also incentivize responsible AI use by shaping how users interact with these tools. For example, Trend Micro has introduced deepfake detection and AI-based threat prevention in its cybersecurity products, helping to identify and stop AI-enabled attacks in real time. 

Companies must also stay ahead of regulatory trends to remain competitive. The EU AI Act and proposed U.S. AI oversight frameworks require greater transparency, risk assessment, and accountability in AI deployment, which can help companies proactively align with emerging regulations. 

Contextual awareness is equally critical. AI models often fail to navigate local norms and geopolitics: Early ChatGPT models produced politically biased outputs; Google Translate mistranslated politically sensitive terms; AI deepfakes fueled misinformation in Asian elections. These examples highlight the need for monitoring, moderation, and user policies adapted to local conditions.

Crucially, mitigating AI misuse can also be commercially advantageous. Companies that invest in safeguards such as content traceability and misuse detection can differentiate themselves as trustworthy providers. They can attract enterprise clients, and reduce legal and reputational risk, gaining a competitive edge even in the absence of strict regulation. 

AI companies must innovate responsibly. By proactively testing models, sharing intelligence across the industry, and implementing traceability safeguards, companies can embed ethics and geopolitical foresight into product design. By protecting societal trust and stability, companies can help generative AI fulfill its full potential as a transformative technology.