Trustworthy AI
发布时间:2026-09-18 | 浏览:2
Artificial intelligence systems have become increasingly prevalent in everyday life and enterprise settings, and they’re now often being used to support human decision-making. These systems have grown increasingly complex and efficient, and AI holds the promise of uncovering valuable insights across a wide range of applications. But broad adoption of AI systems will require humans to trust their output.
When people understand how technology works, and we can assess that it’s safe and reliable, we’re far more inclined to trust it. Many AI systems to date have been black boxes, where data is fed in and results come out. To trust a decision made by an algorithm, we need to know that it is fair, that it’s reliable and can be accounted for, and that it will cause no harm. We need assurances that AI cannot be tampered with and that the system itself is secure. We need to be able to look inside AI systems, to understand the rationale behind the algorithmic outcome, and even ask it questions as to how it came to its decision.
At IBM Research, we’re working on a range of approaches to ensure that AI systems built in the future are fair, robust, explainable, account, and align with the values of the society they’re designed for. We’re ensuring that in the future, AI applications are as fair as they are efficient across their entire lifecycle.
Bringing a common language to AI evaluation News Kim Martineau 23 Jul 2026 AI AI Transparency Fairness, Accountability, Transparency Generative AI
Bringing a common language to AI evaluation
AI Transparency
Fairness, Accountability, Transparency
How the wrong training environment can teach AI models to misbehave Research Peter Hess 09 Jul 2026 AI AI Planning Explainable AI Generative AI
How the wrong training environment can teach AI models to misbehave
Introducing the IBM Granite 4.1 family of models Release Mike Murphy 29 Apr 2026 AI Computer Vision Foundation Models Generative AI Speech Trustworthy AI
Introducing the IBM Granite 4.1 family of models
Computer Vision
Foundation Models
Toward a transparent supply chain for AI News Kim Martineau 26 Mar 2026 AI AI Transparency Data and AI Security
Toward a transparent supply chain for AI
AI Transparency
Data and AI Security
How IBM Granite became a leader in responsible AI Explainer Kim Martineau 13 Feb 2026 AI Generative AI Trustworthy AI Trustworthy Generation
How IBM Granite became a leader in responsible AI
Trustworthy Generation
LLMs have model cards. Now, benchmarks do, too Release Kim Martineau 16 Dec 2025 AI Fairness, Accountability, Transparency Generative AI Trustworthy AI
LLMs have model cards. Now, benchmarks do, too
Fairness, Accountability, Transparency
See more of our work on Trustworthy AI
AI Testing We’re designing tools to help ensure that AI systems are trustworthy, reliable and can optimize business processes.
Adversarial Robustness and Privacy We’re making tools to protect AI and certify its robustness, and helping AI systems adhere to privacy requirements.
Adversarial Robustness and Privacy
Explainable AI We’re creating tools to help AI systems explain why they made the decisions they did.
Fairness, Accountability, Transparency We’re developing technologies to increase the end-to-end transparency and fairness of AI systems.
Fairness, Accountability, Transparency
Trustworthy Generation We’re developing theoretical and algorithmic frameworks for generative AI to accelerate future scientific discoveries.
Trustworthy Generation
Uncertainty Quantification We’re developing ways for AI to communicate when it's unsure of a decision across the AI application development lifecycle.
Uncertainty Quantification
Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting Zhi-yi Chin Pin-Yu Chen et al. 2026 COLM 2026 Conference paper
Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting
Zhi-yi Chin Pin-Yu Chen et al.
From Alignment to Access Control: A Unified View of GenAI Policy Enforcement Nathalie Baracaldo Angel 2026 USENIX Security 2026 Talk
From Alignment to Access Control: A Unified View of GenAI Policy Enforcement
Nathalie Baracaldo Angel
Nathalie Baracaldo Angel
USENIX Security 2026
The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search Rongzhe Wei Peizhi Niu et al. 2026 ICML 2026 Conference paper
The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
Rongzhe Wei Peizhi Niu et al.
ProbeLLM: Automating Principled Diagnosis of LLM Failures Yue Huang Zhengzhe Jiang et al. 2026 ICML 2026 Conference paper
ProbeLLM: Automating Principled Diagnosis of LLM Failures
Yue Huang Zhengzhe Jiang et al.
Position: Agentic Systems Should be General Elron Bandel Asaf Yehudai et al. 2026 ICML 2026 Conference paper
Position: Agentic Systems Should be General
Elron Bandel Asaf Yehudai et al.
Scaling Laws in Model Fine-tuning for Audio DeepFake Detection Xiang Li Pin-Yu Chen et al. 2026 ICML 2026 Conference paper
Scaling Laws in Model Fine-tuning for Audio DeepFake Detection
Xiang Li Pin-Yu Chen et al.
Building trustworthy AI with Watson
Our research is regularly integrated into Watson solutions to make IBM’s AI for business more transparent, explainable, robust, private, and fair.