AI's Confidence Problem: Why Machines Struggle to Say 'I Don't Know' in Healthcare
Some AI models confabulated fake drug names up to 99.6% of the time, revealing a dangerous gap in healthcare AI safety and reliability.
120 articles
Some AI models confabulated fake drug names up to 99.6% of the time, revealing a dangerous gap in healthcare AI safety and reliability.
25 Fields Medallists, including Terence Tao, warn AI's math problem obsession threatens the human insight that keeps mathematics alive.
A new benchmark reveals some AI materials science models get the right answers for the wrong reasons, and Meta, Microsoft, and others are already adopting.
Quantum computers may need just 100,000 qubits to break RSA-2048 encryption, and experts say organizations must migrate to quantum-resistant cryptography.
Lila Sciences is betting on open-ended AI search, a strategy inspired by evolution, to autonomously drive breakthroughs in materials science, energy, and.
China's Supreme Court set the world's first national AI liability standard, making fault-based rules for deepfakes, copyright, and autonomous vehicles.
Insilico Medicine's AI-discovered drug Rentosertib has entered Phase III trials for IPF, a fatal lung disease with no cure and only 3-4 years median.
OpenAI's Navier-Stokes breakthrough faces plagiarism allegations after researchers claim the company may have used their private AI sessions without.
Two mathematicians used AI to make major progress on the Navier-Stokes equations, edging closer to a $1 million Millennium Prize solution.
IFM released six open-source AI models up to 375B parameters and audited its own benchmarks, catching reward hacking that inflated scores by 3.37 points.
AI agents are now running chemistry labs 24/7, with one system autonomously synthesizing 41 novel materials in just 17 days without human intervention.
A new study introduces the MURP framework, pushing healthcare AI beyond accuracy to evaluate transparency, uncertainty, and provenance for regulated.
AI chatbots hallucinate confident wrong answers because they predict likely words, not facts; understanding this single flaw helps users verify outputs.
Canadian scientists used 7-Tesla MRI to map the brain's claustrum, a consciousness hub too thin to image until now, opening new research into human.
AI's real race is now about power plants, not models; NVIDIA's $12.9B Hugging Face deal and Anthropic's 460-megawatt compute pact prove it.
Pathway's BDH-CQ AI model cuts reasoning costs 11 times versus GPT 5.6 Luna using a post-transformer architecture that solves puzzles like humans do.
Anthropic's AI system improved itself on every alignment benchmark tested, outperforming human researchers in six hours at just $4 per hour versus $150.
Puzzles remain AI research's sharpest diagnostic tool, exposing reasoning gaps and generalization failures that predict real-world vulnerabilities before.
MIT's CrysVCD framework slashes material design screening time by 90% by enforcing chemical rules before generation, unlocking faster innovation for chips.
Only 2.6% of companies tie AI to executive incentives, revealing a gap between AI research ambition and the culture, deployment, and accountability needed.
Ten major AI conferences will crowd the 2026 calendar; here is how to pick the right one based on what you actually build.
Major universities are quietly embedding DEI into AI tools and curricula before K-12 schools adopt them, shaping how millions of students will learn.
Apple's M5 Ultra chip lets researchers run frontier AI models locally on 512GB of unified memory, with no cloud fees and full data privacy.
AI may finally crack turbulence, one of physics' hardest problems, helping engineers design better aircraft, power plants, and fusion reactors faster.
AI researcher Oren Etzioni's new glossary exposes how vague AI jargon hides real costs, with inference now accounting for two-thirds of all AI compute.
Finnish researchers built an AI that reads like humans do, but the cognitive profiles it creates to personalize text raise serious privacy concerns.
AI is nearing a shift from answering questions to making discoveries, with biomedical research poised to be transformed by true AI innovation.
AI drug discovery agents can now run entire research workflows autonomously, but experts say safety guardrails and human oversight are equally critical to.
Micron is investing $10 billion in a U.S. memory research lab, betting that AI's future depends on memory innovation, not just faster processors.
Google's new ME-POIs framework helps AI understand cities like locals, boosting visit intent prediction by 81.9% using real movement data.
Patients abandon AI healthcare tools not from tech failures, but distrust in care quality, new research from Binghamton University reveals.
Microsoft's Skala 1.1 brings near-hybrid DFT accuracy to five major chemistry platforms, letting researchers run faster simulations without sacrificing.
AI satellite surveillance meant to monitor nuclear threats may actually increase conflict risk, researchers warn, due to fragility, physics limits, and.
Agentic AI is projected to surge from $19 billion to $206 billion by 2033, as enterprises shift from AI copilots to autonomous systems that complete.
GEN-1.5 robots now learn dexterous tasks from a single 3-second demo, hitting 59% success rates and accelerating AGI timelines faster than expected.
UC Berkeley's agentic AI course has reached nearly 40,000 learners, as universities race to teach autonomous AI agent safety before industry deploys the.
Workday's new AI research team found deleted data lingers in agent memory 20% of the time, revealing why enterprise AI needs smarter design to truly.
With 40+ AI chip makers and no agreed standard, MLPerf benchmarks offer the first fair method to compare AI hardware performance across competing systems.
Quantinuum, NVIDIA, and Pfizer's ADAPT-GQE framework uses generative AI to automate quantum circuit design, accelerating drug discovery on real.
Deep Origin's physics-plus-AI drug discovery framework hit a 31% success rate on a cancer target, nearly 100 times better than pure machine learning.
Cornell Tech's six new AI researchers are tackling machine learning's hardest problems, from trustworthy unlearning benchmarks to efficient reasoning.
AI research agents succeed at only one-third of real science tasks in 2026 benchmarks, exposing a wide gap between vendor hype and actual reliability.
EdUHK's three ICML 2026 papers, one chosen from just 168 oral spotlights, reveal how to make AI admit uncertainty, recommend fairly, and predict honestly.
USC's quantum AI system spots tumors 7% more accurately using six times fewer computing parameters, potentially cutting radiation planning from days to.
Insilico Medicine's Virtual Aging Cell uses multi-agent AI to simulate how cells change over time, a first step toward modeling aging across six.
Google DeepMind's sign language AI brings ASL-to-text dictation to Pixel 11 phones, reaching 70 million Deaf users left behind by voice technology.
USC researchers found audio AI models ignore emotional tone, boosting accuracy from 17% to 65% with fixes that help AI truly listen, not just transcribe.
A Carnegie analysis finds eight structural barriers blocking Pentagon AI adoption, even as AI-assisted targeting already processes 5,000 military targets.
Transformers power every major LLM, but startups are racing to replace them as OpenAI alone prepares to spend $50 billion on computing in 2026.
AI chip deals and a $100M-per-model training cost are reshaping who wins the AI race, as hardware infrastructure becomes the critical bottleneck in 2026.