GPT-5.2, Gemini e Claude Vencem OpenEvidence e UpToDate: o Benchmark da Nature Medicine e a ColisΓ£o com a CFM 2.454
GPT-5.2, Gemini e Claude Vencem OpenEvidence e UpToDate: o Benchmark da Nature Medicine e a ColisΓ£o com a CFM 2.454
Em junho, um estudo da Nature Medicine mostrou que os modelos generalistas β GPT-5.2, Gemini 3.1 Pro e Claude Opus 4.6 β superam as ferramentas clΓnicas dedicadas (OpenEvidence e UpToDate Expert AI) em todos os benchmarks. Dois meses depois entra em vigor a CFM 2.454, que exige rastreabilidade e mediaΓ§Γ£o humana. A ferramenta que funciona melhor Γ© justamente a que mais expΓ΅e o mΓ©dico.
π Publicado em 8 de julho de 2026
Por Que Este Cruzamento Importa
Em 23 de junho de 2026, a Nature Medicine publicou um dado que deveria incomodar todo gestor de saΓΊde: os modelos de IA generalistas β nΓ£o construΓdos especificamente para medicina β superaram as ferramentas clΓnicas dedicadas em todos os benchmarks testados. GPT-5.2, Gemini 3.1 Pro e Claude Opus 4.6 bateram o OpenEvidence e o UpToDate Expert AI em conhecimento, alinhamento e, sobretudo, em perguntas clΓnicas do mundo real.
Sessenta e quatro dias depois, o Brasil coloca em vigor a ResoluΓ§Γ£o CFM nΒΊ 2.454/2026. A pergunta que importa nΓ£o Γ© “a IA vai substituir o mΓ©dico?” β essa Γ© preguiΓ§osa. A pergunta que importa para quem opera no chΓ£o do hospital Γ©: o que acontece quando a ferramenta que funciona melhor Γ© exatamente a que deixa o mΓ©dico mais exposto? Esse Γ© o paradoxo que o intensivista brasileiro vai viver a partir de agosto.
O Que a Nature Medicine Mediu
O estudo comparou trΓͺs modelos de propΓ³sito geral β GPT-5.2 (OpenAI), Gemini 3.1 Pro (Google) e Claude Opus 4.6 (Anthropic) β contra duas ferramentas clΓnicas dedicadas e amplamente adotadas: o OpenEvidence (usado por mais de 750 mil mΓ©dicos verificados) e o UpToDate Expert AI (Wolters Kluwer). A avaliaΓ§Γ£o teve trΓͺs estΓ‘gios: 500 questΓ΅es estruturadas do MedQA (conhecimento), 500 itens do HealthBench (alinhamento) e, o mais importante, um benchmark de consultas clΓnicas reais (RCQ) β 100 perguntas desidentificadas que mΓ©dicos de fato enviaram a um modelo de linguagem em ambiente clΓnico ao vivo, avaliadas por 12 clΓnicos norte-americanos em desenho randomizado e cego, gerando 1.800 anotaΓ§Γ΅es.
Por que o terceiro estΓ‘gio muda tudo
Esse estΓ‘gio Γ© o que separa este trabalho da enxurrada de “IA passou na prova de residΓͺncia”. Prova Γ© ambiente controlado, alternativa AβD, resposta ΓΊnica. O bedside Γ© pergunta mal formulada, contexto incompleto, ambiguidade β o territΓ³rio onde a maioria das ferramentas quebra. E foi exatamente ali que as ferramentas dedicadas perderam por mais.
β O que NΓO se deve concluir do paper
Benchmark nΓ£o Γ© desfecho clΓnico. AcurΓ‘cia em MedQA/HealthBench nΓ£o equivale a reduΓ§Γ£o de mortalidade, tempo de internaΓ§Γ£o ou erro terapΓͺutico. Γ sinal de capacidade, nΓ£o prova de seguranΓ§a em deploy prospectivo. Nenhum desses nΓΊmeros autoriza uso autΓ΄nomo em decisΓ£o terapΓͺutica β a leitura correta Γ© sobre qual ferramenta apoia melhor, nΓ£o sobre substituir o julgamento.
Os NΓΊmeros, Sem Filtro
| Ferramenta | MedQA (%) | HealthBench (/100) | RCQ (mΓ©dia) | Categoria |
|---|---|---|---|---|
| Gemini 3.1 Pro | 97,4 | 79,3 | 3,62 | Generalista |
| GPT-5.2 | 94,2 | 88,0 | 3,54 | Generalista |
| Claude Opus 4.6 | 90,2 | 77,0 | 3,52 | Generalista |
| OpenEvidence | 89,6 | 62,6 | 3,24 | Especializado |
| UpToDate Expert AI | 88,4 | 61,3 | 3,17 | Especializado |
No MedQA a diferenΓ§a Γ© modesta. No HealthBench ela vira abismo: cerca de 26 pontos entre o melhor generalista (GPT-5.2, 88,0) e a melhor ferramenta dedicada (OpenEvidence, 62,6). E na consulta clΓnica real, os generalistas ficaram no topo (Gemini 3,62; GPT 3,54; Claude 3,52) contra 3,24 do OpenEvidence e 3,17 do UpToDate β com as ferramentas dedicadas tendo de 49% a 87% menos chance de receber a nota mais alta do clΓnico do que o Gemini. Detalhe que dΓ³i: nessa etapa, as ferramentas dedicadas performaram de modo comparΓ‘vel ao AI Overview do Google.
A leitura correta
NΓ£o Γ© “generalista Γ© melhor que especializado” em abstrato. Γ que o produto construΓdo e vendido para uso clΓnico nΓ£o sustenta superioridade justamente no tipo de pergunta que mais importa Γ beira do leito β a consulta nΓ£o estruturada que chega no meio de um plantΓ£o. O selo de “ferramenta clΓnica dedicada”, hoje, nΓ£o Γ© garantia de melhor desempenho.
PrecisΓ£o que te blinda entre pares
Cuidado com a manchete “IA generalista vence IA com clearance do FDA”. OpenEvidence e UpToDate Expert AI sΓ£o plataformas de decisΓ£o clΓnica β nΓ£o sΓ£o, elas prΓ³prias, dispositivos mΓ©dicos registrados. O que ganhou clearance do FDA em jun/2026 foi o EchoNext (Pathway Labs), um detector de cardiopatia estrutural em ECG que a OpenEvidence passou a integrar β coisa distinta da ferramenta de resposta que o estudo mediu. Repetir “FDA-cleared” para o produto errado Γ© o tipo de imprecisΓ£o que um revisor atento derruba.
A CFM 2.454 na PrΓ‘tica
Publicada no DOU em 27 de fevereiro de 2026, a ResoluΓ§Γ£o entra em vigor 180 dias depois β em 26 de agosto de 2026. Ela nΓ£o proΓbe IA; ela ancora responsabilidade. TrΓͺs pilares importam para quem trabalha no chΓ£o do hospital:
MediaΓ§Γ£o humana obrigatΓ³ria
Γ vedado delegar Γ IA a comunicaΓ§Γ£o de diagnΓ³sticos, prognΓ³sticos ou condutas terapΓͺuticas. O mΓ©dico revisa e valida antes de a informaΓ§Γ£o chegar ao paciente. A decisΓ£o final Γ© sempre humana.
ClassificaΓ§Γ£o por risco
Sistemas sΓ£o categorizados em baixo, mΓ©dio, alto ou inaceitΓ‘vel (Anexo II), conforme impacto em direitos fundamentais, autonomia do modelo e sensibilidade dos dados. AplicaΓ§Γ΅es de risco inaceitΓ‘vel sΓ£o simplesmente incompatΓveis com a norma β e a obrigaΓ§Γ£o de classificar Γ© da instituiΓ§Γ£o, nΓ£o do fornecedor.
Rastreabilidade no prontuΓ‘rio
O uso de IA como apoio Γ decisΓ£o deve ser registrado, garantindo trilha auditΓ‘vel e seguranΓ§a jurΓdica.
Detalhe que quase ninguΓ©m leu
A ResoluΓ§Γ£o funciona tambΓ©m como escudo: o mΓ©dico nΓ£o responde por falha atribuΓvel exclusivamente ao sistema, desde que demonstre uso diligente, crΓtico e Γ©tico. Ou seja β o registro no prontuΓ‘rio nΓ£o Γ© burocracia, Γ© a sua proteΓ§Γ£o jurΓdica. Quem documenta o uso crΓtico da IA transfere o Γ΄nus da falha sistΓͺmica para onde ele pertence.
O Paradoxo da Ferramenta Dedicada
Junte as duas notΓcias e o desenho fica claro. A ferramenta que performa melhor Γ© o LLM generalista, que nΓ£o foi construΓdo para medicina. A ferramenta com pedigree clΓnico β feita, curada e vendida para o mΓ©dico β Γ© justamente a que a Nature Medicine mostrou perder na consulta real. O intensivista brasileiro tende a usar o GPT-5.2 ou o Claude porque funcionam, mas o faz num vΓ‘cuo de governanΓ§a: o modelo genΓ©rico nΓ£o nasce com trilha de auditoria clΓnica.
β O “pedigree clΓnico” nΓ£o Γ© garantia de superioridade
Confiar cegamente na “ferramenta feita para mΓ©dico” pode significar usar a opΓ§Γ£o que rende menos na pergunta que importa. E usar o modelo generalista superior sem registrar o uso crΓtico pode significar responder sozinho por uma falha que nΓ£o Γ© sua. O ponto de equilΓbrio nΓ£o estΓ‘ em escolher entre performance e conformidade β estΓ‘ em construir conformidade sobre a performance.
SupervisΓ£o Significativa e Rastreabilidade
O conceito mais subestimado da resoluΓ§Γ£o Γ© o de supervisΓ£o significativa. Ele inverte o Γ΄nus da prova da seguranΓ§a: nΓ£o basta afirmar “tem humano no laΓ§o”; Γ© preciso provar que o humano tem condiΓ§Γ΅es reais de revisar β tempo, interface, contexto clΓnico e capacidade de auditar o que a IA propΓ΄s.
E aqui a evidΓͺncia da prΓ³pria Nature Medicine reforΓ§a o ponto: se o gargalo jΓ‘ nΓ£o Γ© a capacidade bruta do modelo, o diferencial passa a ser a camada de governanΓ§a em torno dele. Um chat genΓ©rico brilha na resposta, mas nΓ£o gera log clΓnico auditΓ‘vel por padrΓ£o. Γ exatamente essa lacuna que a CFM 2.454 obriga a fechar.
A Linha do Tempo
Da publicaΓ§Γ£o da norma ao benchmark que a contradiz
Por que a coincidΓͺncia de calendΓ‘rio entre a Nature Medicine e a CFM 2.454 define a janela de decisΓ£o do mΓ©dico.
O Que Fazer na Segunda-Feira
Da teoria regulatΓ³ria Γ conduta de plantΓ£o
Use IA no Bedside com Rigor ClΓnico e Cobertura RegulatΓ³ria
A metodologia AIMED forma mΓ©dicos que usam IA com rigor crΓtico, consciΓͺncia regulatΓ³ria e aplicaΓ§Γ£o clΓnica real β do prompt Γ trilha no prontuΓ‘rio, com framework de implementaΓ§Γ£o e casos aplicados. Prepare-se antes de agosto.
ConheΓ§a o AIMED βConsideraΓ§Γ΅es Finais
O benchmark da Nature Medicine nΓ£o Γ© uma ameaΓ§a ao mΓ©dico brasileiro. Γ um espelho. Ele mostra que a capacidade bruta migrou para o modelo generalista β e, por contraste, ilumina o que continua sendo nosso: o julgamento sob incerteza, a responsabilidade pelo paciente concreto e a capacidade de auditar a prΓ³pria ferramenta.
A ResoluΓ§Γ£o 2.454 nΓ£o freou o futuro; apenas exigiu que o futuro tenha um nome assinado embaixo. E esse nome, por lei e por mΓ©rito, continua sendo o do mΓ©dico.
π‘ Connecting the Dots: a CFM 2.454 amarra a proteΓ§Γ£o jurΓdica do mΓ©dico Γ rastreabilidade do uso crΓtico, enquanto a Nature Medicine mostra que a melhor ferramenta Γ© a que menos oferece essa trilha por padrΓ£o. A vantagem competitiva do mΓ©dico-desenvolvedor estΓ‘ exatamente nessa fenda: construir a camada de rastreabilidade em torno do modelo generalista superior β capturar prompt, versΓ£o, resposta e validaΓ§Γ£o humana no prontuΓ‘rio β Γ© o que transforma a ferramenta mais capaz na ferramenta mais defensΓ‘vel. NΓ£o Γ© escolher entre performance e conformidade; Γ© engenharia de conformidade sobre a performance. Quem dominar isso primeiro define o padrΓ£o para o SUS.
ReferΓͺncias
- General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nature Medicine. 2026 jun 23. DisponΓvel em: https://www.nature.com/articles/s41591-026-04431-5
- Conselho Federal de Medicina. ResoluΓ§Γ£o CFM nΒΊ 2.454, de 11 de fevereiro de 2026. DOU 2026 fev 27; Ed. 39, SeΓ§Γ£o 1, p. 158. DisponΓvel em: https://sistemas.cfm.org.br/normas/arquivos/resolucoes/BR/2026/2454_2026.pdf
- Conselho Federal de Medicina. CFM normatiza uso da IA na medicina. Portal MΓ©dico. 2026. DisponΓvel em: https://portal.cfm.org.br/noticias/cfm-normatiza-uso-da-ia-na-medicina/
GPT-5.2, Gemini and Claude Beat OpenEvidence and UpToDate: The Nature Medicine Benchmark and Its Collision with CFM 2.454
In June, a Nature Medicine study showed generalist models β GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 β outperform dedicated clinical tools (OpenEvidence and UpToDate Expert AI) across every benchmark. Two months later, CFM 2.454 takes effect, demanding traceability and human mediation. The tool that works best is precisely the one that leaves the physician most exposed.
π Published July 8, 2026
Why This Cross-Reading Matters
On June 23, 2026, Nature Medicine published a finding that should unsettle every health administrator: generalist AI models β not purpose-built for medicine β outperformed dedicated clinical tools across every benchmark tested. GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 beat OpenEvidence and UpToDate Expert AI in knowledge, alignment, and above all in real-world clinical queries.
Sixty-four days later, Brazil enacts CFM Resolution 2.454/2026. The question that matters is not “will AI replace the physician?” β that one is lazy. The question that matters for those on the hospital floor is: what happens when the tool that works best is exactly the one that leaves the physician most exposed? That is the paradox the Brazilian intensivist will face from August on.
What Nature Medicine Measured
The study compared three general-purpose models β GPT-5.2 (OpenAI), Gemini 3.1 Pro (Google), and Claude Opus 4.6 (Anthropic) β against two widely adopted dedicated clinical tools: OpenEvidence (used by over 750,000 verified physicians) and UpToDate Expert AI (Wolters Kluwer). The evaluation had three stages: 500 structured MedQA questions (knowledge), 500 HealthBench items (alignment), and most importantly a real clinical queries (RCQ) benchmark β 100 de-identified questions physicians actually sent to a language model in a live clinical setting, reviewed by 12 US clinicians in a randomized, blinded design producing 1,800 annotations.
Why the third stage changes everything
That stage sets this apart from the flood of “AI passes the board exam” headlines. Exams are controlled: AβD options, single answer. The bedside is a poorly framed question, incomplete context, ambiguity β the terrain where most tools break. And that is exactly where the dedicated tools lost by the widest margin.
β What NOT to conclude from the paper
A benchmark is not a clinical outcome. Accuracy on MedQA/HealthBench does not equal reduced mortality, length of stay, or therapeutic error. It signals capability, not prospective deployment safety. None of these numbers authorize autonomous use in therapeutic decisions β the correct reading is about which tool supports better, not about replacing judgment.
The Numbers, Unfiltered
| Tool | MedQA (%) | HealthBench (/100) | RCQ (mean) | Category |
|---|---|---|---|---|
| Gemini 3.1 Pro | 97.4 | 79.3 | 3.62 | Generalist |
| GPT-5.2 | 94.2 | 88.0 | 3.54 | Generalist |
| Claude Opus 4.6 | 90.2 | 77.0 | 3.52 | Generalist |
| OpenEvidence | 89.6 | 62.6 | 3.24 | Dedicated |
| UpToDate Expert AI | 88.4 | 61.3 | 3.17 | Dedicated |
On MedQA the gap is modest. On HealthBench it becomes a chasm: about 26 points between the best generalist (GPT-5.2, 88.0) and the best dedicated tool (OpenEvidence, 62.6). And on real clinical queries, the generalists clustered at the top (Gemini 3.62; GPT 3.54; Claude 3.52) versus 3.24 for OpenEvidence and 3.17 for UpToDate β with the dedicated tools having 49% to 87% lower odds of receiving the highest clinician rating than Gemini. A detail that stings: on this stage, the dedicated tools performed comparably to Google’s AI Overview.
The correct reading
It is not “generalist beats specialized” in the abstract. It is that the product built and sold for clinical use fails to hold superiority precisely on the question type that matters most at the bedside β the unstructured query that arrives mid-shift. The “dedicated clinical tool” label, today, is no guarantee of better performance.
Precision that shields you among peers
Beware the headline “generalist AI beats FDA-cleared AI.” OpenEvidence and UpToDate Expert AI are clinical decision platforms β they are not themselves registered medical devices. What received FDA clearance in June 2026 was EchoNext (Pathway Labs), a structural-heart-disease detector on ECG that OpenEvidence began to integrate β a different thing from the answer tool the study measured. Repeating “FDA-cleared” about the wrong product is the kind of imprecision an attentive reviewer takes down.
CFM 2.454 in Practice
Published in the Official Gazette on February 27, 2026, the Resolution takes effect 180 days later β on August 26, 2026. It does not ban AI; it anchors responsibility. Three pillars matter for those working on the hospital floor:
Mandatory human mediation
Delegating the communication of diagnoses, prognoses, or therapeutic decisions to AI is prohibited. The physician reviews and validates before information reaches the patient. The final decision is always human.
Risk classification
Systems are categorized as low, medium, high, or unacceptable (Annex II), by impact on fundamental rights, model autonomy, and data sensitivity. Unacceptable-risk applications are simply incompatible with the norm β and the duty to classify falls on the institution, not the vendor.
Traceability in the medical record
AI use as decision support must be recorded, ensuring an auditable trail and legal security.
The detail almost no one read
The Resolution also works as a shield: the physician is not liable for a failure attributable exclusively to the system β provided diligent, critical, and ethical use is demonstrated. In other words, the medical-record entry is not bureaucracy; it is your legal protection. Documenting critical AI use shifts the burden of systemic failure to where it belongs.
The Dedicated-Tool Paradox
Put the two stories together and the picture is clear. The best-performing tool is the generalist LLM, which was not built for medicine. The tool with clinical pedigree β made, curated, and sold to physicians β is precisely the one Nature Medicine showed loses on real queries. Brazilian intensivists tend to use GPT-5.2 or Claude because they work, but do so in a governance vacuum: the generic model is not born with a clinical audit trail.
β “Clinical pedigree” is no guarantee of superiority
Blindly trusting the “tool made for physicians” may mean using the option that scores lower on the question that matters. And using the superior generalist model without recording critical use may mean answering alone for a failure that is not yours. The balance is not in choosing between performance and compliance β it is in building compliance on top of performance.
Significant Supervision and Traceability
The most underestimated concept in the resolution is significant supervision. It reverses the burden of proof for safety: it is not enough to claim “there’s a human in the loop”; it must be proven that the human has real conditions to review β time, interface, clinical context, and the ability to audit what the AI proposed.
And here Nature Medicine’s own evidence reinforces the point: if the bottleneck is no longer the model’s raw capability, the differentiator becomes the governance layer around it. A generic chat shines in the answer but generates no auditable clinical log by default. That is exactly the gap CFM 2.454 forces to close.
The Timeline
From the norm’s publication to the benchmark that contradicts it
Why the calendar coincidence between Nature Medicine and CFM 2.454 defines the physician’s decision window.
What to Do on Monday
From regulatory theory to bedside conduct
Use AI at the Bedside with Clinical Rigor and Regulatory Coverage
The AIMED methodology develops physicians who use AI with critical rigor, regulatory awareness, and real clinical application β from prompt to record trail, with an implementation framework and applied cases. Get ready before August.
Discover AIMED βFinal Considerations
The Nature Medicine benchmark is not a threat to the Brazilian physician. It is a mirror. It shows that raw capability has migrated to the generalist model β and, by contrast, illuminates what remains ours: judgment under uncertainty, responsibility for the concrete patient, and the ability to audit the tool itself.
Resolution 2.454 did not stop the future; it merely required that the future carry a signed name beneath it. And that name, by law and by merit, remains the physician’s.
π‘ Connecting the Dots: CFM 2.454 ties the physician’s legal protection to traceability of critical use, while Nature Medicine shows the best tool is the one that offers that trail least by default. The physician-developer’s competitive edge lies exactly in that gap: building the traceability layer around the superior generalist model β capturing prompt, version, response, and human validation in the record β is what turns the most capable tool into the most defensible one. It is not choosing between performance and compliance; it is compliance engineering on top of performance. Whoever masters this first sets the standard for the public health system.
References
- General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nature Medicine. 2026 Jun 23. Available at: https://www.nature.com/articles/s41591-026-04431-5
- Federal Council of Medicine (CFM). Resolution CFM No. 2,454, of February 11, 2026. Official Gazette 2026 Feb 27; Ed. 39, Sec. 1, p. 158. Available at: https://sistemas.cfm.org.br/normas/arquivos/resolucoes/BR/2026/2454_2026.pdf
- Federal Council of Medicine. CFM regulates AI use in medicine. Portal MΓ©dico. 2026. Available at: https://portal.cfm.org.br/noticias/cfm-normatiza-uso-da-ia-na-medicina/
