Logo
FrontierNews.ai

OpenAI's ChatGPT Health Launches With Major Safety Questions Unanswered

OpenAI has launched ChatGPT Health, a new feature that connects to your medical records and Apple Watch data, but independent research reveals it misses genuine emergencies more than half the time. The company introduced the tool last month to select US adults, allowing the AI to help explain doctor visit notes, interpret lab results, and prepare questions for appointments. However, the rollout has sparked alarm among medical researchers who question whether the system is ready for real-world use.

What Makes ChatGPT Health Different From Regular ChatGPT?

ChatGPT Health builds on the same underlying technology as the standard ChatGPT, but with a critical addition: it can access your personal health data. With your permission, the AI can link to your medical records and download health metrics from Apple wearables like the Apple Watch. OpenAI initially created it as a separate tab in ChatGPT's interface, but after observing users naturally asking health questions during regular conversations, the company seamlessly integrated it into the main platform.

The feature represents OpenAI's attempt to address a real gap in healthcare access. "The reality is, we're living in a time when not everyone has access to quality care or care in a timely manner," an OpenAI executive told reporters. "The average doctor appointment in the United States is less than 15 minutes. You're often left on your own to sort through a lot of fragmented information".

Why Are Researchers Concerned About Safety?

The concern centers on a fundamental mismatch: OpenAI says ChatGPT Health is "not designed to replace the care and judgment of qualified medical professionals," yet that's exactly how millions of people will use it. More than 300 million people consult ChatGPT for health-related questions every week, and over 40 million turn to it with healthcare questions daily.

Multiple independent studies have documented serious performance gaps. In a structured test of triage recommendations, researchers at Mount Sinai Health System in New York found that ChatGPT Health under-triaged 52% of genuine emergencies, directing patients with diabetic ketoacidosis or impending respiratory failure to wait 24 to 48 hours for evaluation instead of going to the emergency department. A separate preprint study found that ChatGPT Health agrees with nurse and physician triage decisions only 50% of the time.

A 52% under-triage rate for real emergencies is not a minor flaw. The consequences are asymmetrical and dangerous: while over-triaging a minor ailment might stress an already overburdened medical system, under-triaging a genuine emergency can be catastrophic. Broader research on AI chatbots shows the problem extends beyond ChatGPT Health. One recent study found that chatbots produced "problematic" responses to medical queries between 20% and 50% of the time, with "unsafe" responses ranging from 5% to 13%. An audit published in the British Medical Journal found roughly 50% of chatbot responses to health questions contained problems, with one-fifth rated as "highly problematic" and potentially harmful.

How Is OpenAI Addressing These Safety Concerns?

OpenAI has developed an assessment framework called Health Bench to evaluate ChatGPT Health's performance. Created in collaboration with 262 physicians from 60 countries across 26 medical specialties, Health Bench periodically audits thousands of conversations to identify and repair weaknesses in the platform. However, no data from these audits have been published yet, and the system remains a work in progress rather than a finished product.

The challenge lies in distinguishing between lower-risk functions and clinical judgment. ChatGPT Health can accurately translate clinical terminology, organize medical records, plot and analyze lab results chronologically, and help users prepare questions for their clinician. But the moment it interprets symptoms and offers advice, it becomes a functioning triage system, regardless of OpenAI's disclaimers. The AI itself shows no evidence of understanding this critical boundary.

Steps to Safely Use ChatGPT Health if You Choose To

  • Treat It as a Research Tool Only: Use ChatGPT Health to understand medical terminology and organize your health information, but never rely on it for diagnosis, triage, or treatment decisions without consulting a qualified healthcare provider.
  • Verify All Recommendations With Your Doctor: If ChatGPT Health suggests you need urgent care or recommends a specific course of action, contact your physician or call a nurse hotline before following the advice, especially for symptoms that could indicate an emergency.
  • Recognize Its Limitations: Remember that ChatGPT Health has only a 50% agreement rate with professional triage decisions, meaning it misses genuine emergencies roughly half the time and may also over-alarm for minor issues.
  • Use It for Administrative Tasks: Focus on using the tool to prepare questions for appointments, understand lab results you've already discussed with your doctor, or organize medical records rather than seeking initial medical guidance.

The tension at the heart of ChatGPT Health reflects a broader pattern in AI development. Industry leaders, including OpenAI's Sam Altman, Anthropic's Dario Amodei, and Elon Musk, called for a slowdown in advanced AI development in September 2026, emphasizing the need for stronger safeguards and independent oversight. Yet the same principle that applies to general AI development applies equally to healthcare: safety checks must keep pace with capability. According to experts quoted in the research, that isn't happening.

The comparison to automotive safety is instructive. "Consider the consequences if Ford rolled out a new truck before proving it was safe to drive," one researcher noted. "We wouldn't accept a half-baked product or the manufacturer's reassurance that it was a 'work in progress.' We should demand the same accountability" for AI health tools.

For now, ChatGPT Health remains in limited rollout to select US adults, but the tool's quiet launch and rapid integration into the main ChatGPT interface suggest OpenAI plans broader expansion. The question facing regulators, healthcare providers, and users is whether the current safety validation framework is sufficient before that happens.