
A IA Que Acha a Criança Grave na Fila do Pronto-Socorro — e a Métrica Que Faz Todo Modelo Parecer Melhor do Que É
A IA Que Acha a Criança Grave na Fila do Pronto-Socorro — e a Métrica Que Faz Todo Modelo Parecer Melhor do Que É
Um modelo treinado em 886 mil atendimentos pediátricos identifica, só com dados da triagem, 88% das crianças que vão precisar de intervenção de terapia intensiva — e triplicaria a fração de pacientes de risco intermediário avaliados a tempo. É o tipo de número que rende manchete. Mas há um dado que quase nenhum artigo brasileiro traduz: a métrica que consagrou esses modelos é cega para o que mais importa na triagem — a raridade do evento. Entender por que a AUROC mente em desfecho raro, e o que colocar no lugar dela, é o que separa quem consome escore preditivo de quem sabe julgá-lo.
📅 Publicado em 27 de julho de 2026
Navegue pelo Artigo
Da fila do pronto-socorro à classe de risco regulatória — por que escolher a métrica é escolher o desfecho clínico
- 1. Por Que Isso Importa Para Quem Trabalha na Ponta
- 2. O Que o Estudo Entregou
- 3. Os Números, Sem Filtro
- 4. A Armadilha: A Métrica Que Não Sente a Raridade
- 5. Escolher a Métrica É um Ato Clínico
- 6. A Camada Regulatória — CFM 2.454
- 7. A Linha do Tempo
- 8. O Que Fazer na Segunda-Feira
- 9. Considerações Finais
- 10. Referências
Para quem não é da área — o essencial em 60 segundos
Este artigo funciona em duas camadas. A técnica, para quem trabalha em pronto-socorro ou constrói modelos; e a geral, que só exige curiosidade. Toda ideia central aparece duas vezes: uma com exemplo de hospital, outra com exemplo do dia a dia. Se a primeira não fizer sentido, pule para a segunda — não vai perder nada.
Triagem é o que acontece quando você chega a um pronto-socorro: alguém decide quem é atendido primeiro. Não é ordem de chegada, é ordem de gravidade. No Brasil usamos o Protocolo de Manchester, que separa os pacientes em cinco cores — vermelho é agora, amarelo é logo, verde pode esperar. Nos Estados Unidos usam o ESI, que faz o mesmo com números de 1 a 5.
As siglas de suporte respiratório que vão aparecer são apenas graus de ajuda para respirar, do mais leve ao mais pesado: HFNC (oxigênio em alto fluxo pelo nariz) e nebulização contínua são os leves; CPAP e BiPAP são máscaras que empurram ar com pressão; intubação é o tubo na traqueia, com o aparelho respirando pelo paciente. Vasopressor é a medicação que sustenta a pressão arterial quando o corpo não consegue mais.
Prevalência é a fração de pessoas em quem algo acontece. “Prevalência de 3%” significa que, de cada 100 atendimentos, 3 terminam naquele desfecho. É a palavra mais importante do artigo — guarde essa.
Por Que Isso Importa Para Quem Trabalha na Ponta
Toda triagem pediátrica tem um purgatório. No sistema americano é o ESI nível 3; no Manchester, que é o nosso, é o amarelo. É a faixa onde a criança não está óbvia o suficiente para virar vermelho e não está bem o suficiente para virar verde. É onde mora a bronquiolite que vai cansar em quatro horas, a desidratação que ainda não fechou o tempo de enchimento capilar, a sepse precoce que ainda tem febre e nada mais.
O problema da classificação de risco por níveis não é que ela erre — é que ela satura. Um sistema de cinco níveis tem, por construção, cinco graus de resolução; e a maior parte do movimento acontece dentro de um único nível. O enfermeiro classificador sabe disso. O plantonista sabe disso. E é exatamente essa a lacuna que um modelo preditivo se propõe a preencher: não substituir o Manchester, mas ordenar o que está empilhado dentro do amarelo.
Um estudo publicado em 20 de julho de 2026 fez precisamente isso, em escala. E a resposta técnica que ele deu é boa. Mas o que torna esse artigo digno de leitura crítica não é o modelo — é a métrica que os autores escolheram para reportá-lo, e o motivo de essa escolha ser, ela própria, uma decisão clínica.
O Que o Estudo Entregou
Coorte retrospectiva de 2016 a 2024, em um pronto-socorro pediátrico acadêmico urbano de grande porte nos Estados Unidos, com cerca de 85 mil atendimentos anuais. 886.183 atendimentos. O desfecho ocorreu em 26.721 visitas, ou 3,0% do total.
O que exatamente conta como desfecho — leia antes de comparar com a sua unidade
Os autores definiram “necessidade de intervenção de suporte avançado” como uso de medicação de suporte à vida ou suporte respiratório dentro de 48 horas da chegada ao pronto-socorro. Medicação: infusão contínua de vasopressor, insulina ou terbutalina. Suporte respiratório: intubação, BiPAP, CPAP, heliox, cateter nasal de alto fluxo (HFNC) e pelo menos 2 horas de nebulização contínua de salbutamol. Conta também o que foi feito na enfermaria e o que foi feito em retorno ao pronto-socorro dentro de 48 horas.
Isso é mais largo do que soa. HFNC e salbutamol contínuo por duas horas, no Brasil, acontecem rotineiramente na sala de observação e na enfermaria — não são, para nós, “terapia intensiva”. Boa parte dos 3% é asma e bronquiolite recebendo tratamento padrão, não criança à beira do colapso. Ao ler o número, ajuste a expectativa: o modelo prevê necessidade de escalonamento de suporte, não iminência de parada.
Aqui os autores fizeram a coisa certa e merecem crédito: rodaram um desfecho secundário mais estrito, só com intubação, BiPAP, CPAP e heliox — excluindo HFNC e salbutamol contínuo — e o desempenho se manteve. Também testaram uma janela mais curta, de 8 horas, com resultado semelhante. São duas análises de sensibilidade que a maioria dos artigos não faz.
Os autores treinaram seis algoritmos usando exclusivamente informação disponível no momento da triagem. Nada de exame laboratorial, nada de evolução, nada de reavaliação — só o que o enfermeiro classificador tem em mãos nos primeiros minutos. A rede neural teve o melhor desempenho.
O número que interessa não é o do modelo
Os autores fizeram uma simulação contrafactual: o que aconteceria com o tempo até a avaliação médica se o escore fosse usado junto com o ESI, e não no lugar dele? Na faixa ESI 3 — o purgatório —, a proporção de pacientes de suporte avançado avaliados em tempo hábil saltaria de 23,3% para 75,0%, e a mediana de tempo até o pediatra cairia de 34 para 10 minutos. No ESI 2, de 48,7% para 87,1% (19 → 12 min). No ESI 4, de 10,3% para 60,7% (62 → 7 min).
Repare no desenho: o modelo não reclassifica ninguém. Ele reordena dentro da classe. É uma escolha de arquitetura que preserva o instrumento validado que a equipe já usa e adiciona resolução onde o instrumento é cego. Vale mais como lição de engenharia clínica do que a AUC de qualquer um dos seis algoritmos.
Um exemplo concreto: o que “prever” significa aqui
São 21h de um sábado de inverno. A recepção tem 40 crianças esperando, 26 delas classificadas como amarelo. O Manchester já fez o trabalho dele: separou os 3 vermelhos e os 11 verdes. Sobram 26 amarelos indistinguíveis entre si na tela, ordenados por ordem de chegada.
O modelo não muda a cor de ninguém. Ele reordena os 26 — e coloca no topo a lactente de 5 meses com frequência respiratória no percentil alto para a idade e saturação de 93%, que chegou 40 minutos depois de um adolescente com dor abdominal. Nenhuma dessas informações é nova. Todas estavam na ficha de triagem. O modelo só fez a aritmética que ninguém tem tempo de fazer às 21h de sábado.
⚠ O contrafactual é o elo mais fraco — e vale entender por quê
A regra da simulação é esta: a cada paciente que de fato recebeu suporte avançado, atribui-se o horário de atendimento mais cedo entre todos os pacientes não críticos do mesmo ESI que esperavam simultaneamente. “Avaliado em tempo hábil” significa ter sido visto antes de todos os não críticos daquele mesmo nível.
Repare no que a regra não modela: os falsos-positivos. Com VPP de 32%, a cada 1.000 atendimentos o modelo levanta cerca de 82 sinalizações e apenas 26 são reais — as outras 56 são crianças que também seriam empurradas para a frente da fila, disputando exatamente as mesmas vagas de prioridade. A fila tem capacidade finita: só existe um “próximo pediatra disponível”. A simulação concede o benefício aos verdadeiros-positivos sem cobrar deles a competição dos falsos.
Há uma ironia produtiva aqui, e ela não está na seção de limitações do artigo: o mesmo trabalho que reporta honestamente um VPP de 32% roda um contrafactual que se comporta, na prática, como se o VPP fosse muito maior. Isso não invalida o estudo — mas transforma o 23,3% → 75,0% no teto do que a informação poderia comprar, não no que a implementação entregaria. Some a isso o que os próprios autores admitem: centro único, sem validação prospectiva, com necessidade declarada de recalibração para outros serviços.
⚠ O bloqueio de transferência para o Brasil que quase ninguém vai notar
Está numa única frase da seção de limitações: “nossos modelos dependem de processamento de linguagem natural das narrativas de enfermagem”. O modelo não roda sobre sinais vitais estruturados — ele lê o texto livre que o enfermeiro escreve na triagem.
Isso muda tudo para nós. Primeiro, é NLP treinado em inglês, sobre convenções de registro de enfermagem americanas. Segundo, e mais grave: a classificação de risco brasileira pelo Manchester é fortemente estruturada em discriminadores, e a qualidade e o volume do texto livre variam enormemente entre serviços — em muitos, é uma linha. A variável que mais carrega sinal no modelo original é justamente a que menos existe no fluxo brasileiro. Não é caso de recalibrar: é caso de retreinar sobre outra base de features.
Os Números, Sem Filtro
Antes da tabela, cinco conceitos. Eles parecem áridos, mas cada um responde a uma pergunta que você já faz de cabeça no plantão — só que sem o nome técnico. Vale ler mesmo se estatística não é o seu terreno: o argumento inteiro deste artigo cabe aqui.
Uma nota sobre o “IC95%” que aparece na tabela: é o intervalo de confiança de 95% — a faixa dentro da qual o valor verdadeiro provavelmente está. Faixa estreita, como o 0,59–0,61 do estudo, significa estimativa precisa, efeito de a amostra ser enorme.
O tradutor de métricas (comece pelo limiar — tudo depende dele)
Limiar — “a partir de que ponto eu chamo?” Um modelo não devolve “sim” ou “não”. Devolve um número contínuo, tipo 0,17 ou 0,64. Alguém precisa decidir a partir de qual valor aquilo vira um alerta na tela. Esse ponto de corte é o limiar.
No hospital: a saturação em que você decide chamar o plantonista é um limiar. Se você sobe de 92% para 94%, chama mais cedo e mais vezes: pega quase todas as crianças que iam piorar, e chama muita gente à toa. Se desce para 88%, chama menos e quase sempre com razão — mas deixa passar as que estavam começando a afundar.
Fora dele: é o ajuste de rigor do seu filtro de spam. Apertado demais, e-mail de verdade cai na lixeira. Frouxo demais, propaganda entope a caixa de entrada. Você não consegue os dois — e nenhum ajuste do filtro resolve, porque o problema não está no filtro, está em ter que escolher um ponto na mesma escala.
Sensibilidade e valor preditivo positivo são exatamente esse balanço, e mexer no limiar troca um pelo outro. Não existe ajuste que melhore os dois ao mesmo tempo: é a mesma corda, puxada de pontas opostas. Guarde isso — é a peça que sustenta a conclusão deste artigo.
Sensibilidade — “de quem ia precisar de UTI, quantos o modelo pega?” Sensibilidade 88% quer dizer que, de 100 crianças que vão receber intervenção de terapia intensiva, o modelo acende o alerta em 88. As outras 12 passam. É a métrica que dói errar.
Valor preditivo positivo (VPP) — “quando ele me chama, quantas vezes é de verdade?” VPP 32% quer dizer que, de cada 100 alertas disparados, 32 são crianças que realmente vão precisar. As outras 68 não. Esse é o número que decide se a equipe vai continuar olhando para o alerta depois da terceira semana.
AUROC — “ele sabe dizer qual das duas está mais grave?” Sorteie duas crianças, uma que vai complicar e outra que não. A AUROC é a probabilidade de o modelo dar a nota mais alta para a que complica. É uma prova de ordenação entre pares. E aqui está a chave que este artigo inteiro persegue: essa prova não muda se o evento for comum ou raríssimo. A AUROC é matematicamente insensível à prevalência.
Average Precision (AP) — “das vezes que ele me chamou, quantas valeram a pena?” É o VPP médio, calculado percorrendo todos os limiares possíveis de uma vez só. Em vez de te dar o VPP de um ponto de corte específico, te dá o comportamento do modelo inteiro. E, diferente da AUROC, a AP sente a raridade do evento.
A “linha de base” de uma métrica é a nota que um modelo inútil tira — aquele que sorteia no cara ou coroa. Serve de régua: sem saber a nota do inútil, você não sabe se a nota do bom é boa. E aqui está a diferença que decide tudo, demonstrada por Saito e Rehmsmeier:
A linha de base da AUROC é sempre 0,50, não importa se o desfecho acontece em metade dos pacientes ou em um a cada mil. A linha de base da AP é a própria prevalência — os autores demonstram que ela vale exatamente P/(P+N), a fração de casos positivos no total. Em português de plantão: num desfecho que ocorre em 3% das crianças, chutar no cara ou coroa rende AP de 0,03. Não 0,50 — 0,03.
É por isso que a AP de 0,60 do estudo significa vinte vezes melhor que o acaso. Se fosse AUROC de 0,60, seria pouco acima de cara ou coroa. Mesmo número, leituras opostas — porque as réguas são diferentes.
| Métrica | Valor reportado | Linha de base | Fonte | Leitura |
|---|---|---|---|---|
| Volume da coorte | 886.183 atendimentos (2016–2024) | — | Ha et al. · Hosp Pediatr · 2026 | ↑ robusto |
| Prevalência do desfecho | 26.721 visitas · 3,0% | — | Ha et al. · Hosp Pediatr · 2026 | ⚠ evento raro |
| Average Precision (rede neural) | 0,60 (IC95% 0,59–0,61) | 0,03 | Ha et al. · Hosp Pediatr · 2026 | ↑ 20× o acaso |
| Sensibilidade | 88% (87–89%) | — | Ha et al. · Hosp Pediatr · 2026 | ↑ bom |
| Valor preditivo positivo | 32% (31–32%) | 3% | Ha et al. · Hosp Pediatr · 2026 | ↓ 2 em 3 alertas são falsos |
| Especificidade (derivada, não reportada) | ≈ 94% | — | Cálculo próprio a partir de sens. e VPP | → ver box abaixo |
| ESI 3 avaliados em tempo hábil | 23,3% → 75,0% | — | Ha et al. · análise contrafactual | ↑ teto, não piso |
A conta que reorganiza a tabela inteira
A especificidade não foi reportada no resumo do estudo, mas dá para derivá-la. Com prevalência de 3%, sensibilidade de 88% e VPP de 32%, a cada 1.000 atendimentos: 30 crianças vão precisar de intervenção; o modelo acerta 26 delas; para chegar a um VPP de 32% ele precisa disparar cerca de 82 alertas no total; logo, 56 são falsos positivos. Sobram 914 não-eventos corretamente silenciados de um total de 970. Especificidade ≈ 94%.
Agora leia as duas frases seguintes, que descrevem exatamente o mesmo classificador: “especificidade de 94%” e “dois de cada três alertas são falsos”. A primeira passa em qualquer comitê. A segunda descreve a experiência real da enfermagem no terceiro mês de uso. Não há contradição — há duas perguntas diferentes. Especificidade pergunta o que acontece com quem está bem; VPP pergunta o que acontece com quem foi chamado. Em evento raro, os denominadores são de tamanhos brutalmente diferentes, e por isso as duas respostas divergem tanto.
Se você já passou por um aeroporto, já viveu isso. O detector de metal pega praticamente toda arma que passa por ele — sensibilidade altíssima, e é assim que tem que ser. Ele também fica calado para a esmagadora maioria dos passageiros — especificidade altíssima. E ainda assim, quando ele apita, quase nunca é arma: é fivela de cinto, chave, moeda esquecida no bolso. As três coisas são verdadeiras ao mesmo tempo, e só parecem contraditórias até você notar que arma é raríssima entre passageiros. É exatamente a estrutura do problema deste artigo — e a razão pela qual ninguém desliga o detector de aeroporto, mas todo mundo desliga o alerta do hospital: no aeroporto, o custo do falso alarme é passar o bastão em alguém por dez segundos. No hospital, quando o custo do falso alarme é alto, o alerta morre.
A Armadilha: A Métrica Que Não Sente a Raridade
Aqui está o dado que quase nenhum divulgador brasileiro traduz: a AUROC não muda quando a prevalência muda. Isso não é uma limitação prática nem um viés de amostragem — é uma propriedade matemática. A AUROC é construída sobre sensibilidade e especificidade, e ambas são calculadas dentro de cada grupo: sensibilidade só olha os doentes, especificidade só olha os sadios. Nenhuma das duas sabe a proporção entre os grupos.
A consequência é desconfortável. O mesmo valor de AUROC descreve um modelo com VPP de 90% e um modelo com VPP de 10%, a depender apenas de quão raro é o desfecho. Se você lê “AUROC 0,95” sem saber a prevalência, você não sabe absolutamente nada sobre como o alerta vai se comportar no plantão.
Por que isso é epidêmico na literatura pediátrica
Porque quase todo desfecho grave em pediatria é raro. Parada cardiorrespiratória, sepse, necessidade de ventilação, óbito — todos com prevalência de um dígito. É precisamente o regime em que a AUROC é mais generosa e menos informativa. Um modelo medíocre num desfecho de 2% de prevalência facilmente produz uma AUROC de 0,90, porque é fácil separar a maioria esmagadora de sadios dos poucos doentes; o difícil é acertar qual dos poucos.
Essa crítica não é minha e nem é nova. Davis e Goadrich demonstraram formalmente, em 2006, que a curva precisão-recall é mais informativa em bases desbalanceadas — jargão para “quase todo mundo é negativo e pouquíssimos são positivos”, que é a descrição exata de qualquer desfecho grave em pediatria —, e que otimizar a área sob a ROC não garante otimizar a área sob a curva precisão-recall. Saito e Rehmsmeier, em 2015, foram diretos ao ponto no próprio resumo: a interpretabilidade visual da ROC em dados desbalanceados “pode ser enganosa quanto às conclusões sobre a confiabilidade do desempenho de classificação, devido a uma interpretação intuitiva porém errada da especificidade”. Vinte anos de metodologia estabelecida, quase 4.700 citações, e a literatura clínica pediátrica ainda reporta AUROC em desfecho de 3%.
⚠ Onde meu próprio argumento tem limite: prevalência não é a mesma coisa que espectro
Preciso ser preciso, porque um leitor atento vai encontrar isto sozinho. Dizer que “a AUROC é insensível à prevalência” é uma afirmação sobre a construção da métrica: sensibilidade e especificidade são razões calculadas dentro de cada grupo, e por isso a linha de base da ROC não se move quando a proporção entre os grupos muda.
Isso não significa que a AUROC medida seja a mesma em qualquer serviço. Se a nova população tiver casos mais graves, mais leves, ou uma mistura diferente, as notas que o modelo distribui se espalham de outro jeito — e a AUROC medida na prática muda junto. É o efeito de espectro: não mudou a proporção de doentes, mudou o tipo de doente. São dois problemas distintos que frequentemente aparecem juntos e são confundidos: invariância à prevalência é propriedade matemática da métrica; efeito de espectro é mudança da população.
O estudo que discutimos ilustra bem o segundo problema: a coorte é 61,8% de crianças negras não hispânicas, 23,2% hispânicas ou latinas, com 74% cobertas por seguro público, num único hospital de Washington. É uma população real e bem descrita — e é outra população.
A consequência prática, porém, é a mesma nos dois casos, e reforça o argumento em vez de enfraquecê-lo: a AUROC publicada não permite antecipar quantos alertas falsos a sua equipe vai receber. No primeiro caso porque a métrica não carrega essa informação; no segundo porque a população não é a mesma. Em qualquer dos dois, o número que você precisa é o VPP calculado na sua prevalência.
⚠ O risco não é o falso positivo isolado — é a morte do alerta
Dois de cada três alertas falsos não são um inconveniente estatístico. São o mecanismo do alarm fatigue, e o alarm fatigue tem desfecho clínico documentado: o alerta que ninguém mais olha protege menos que nenhum alerta, porque consome atenção e produz falsa sensação de cobertura. Um sistema de triagem preditiva com VPP baixo e apresentação errada não degrada graciosamente — ele é desligado mentalmente pela equipe em semanas, e continua ligado na tela por meses.
Um exemplo concreto: o alarme de saturação que você já silenciou hoje
Você conhece essa métrica sem saber o nome dela. O oxímetro do leito 4 apita a cada vinte minutos. Na esmagadora maioria das vezes é o sensor deslocado, é a criança que mexeu a mão, é artefato. A especificidade daquele alarme é altíssima — ele fica calado a maior parte do tempo, em relação a todos os minutos em que a criança está bem. O valor preditivo positivo dele é baixíssimo — quando ele grita, quase nunca é dessaturação real.
Fora do hospital, o mesmo: pense no alarme de carro que dispara na rua às três da manhã. Ninguém levanta. Ninguém olha pela janela. Não é porque o alarme nunca acerta — é porque quase sempre erra, e o cérebro humano aprende isso em poucas semanas. Um alarme de carro tem sensibilidade excelente para arrombamento e valor preditivo positivo desprezível, e é por isso que ele virou ruído urbano em vez de sistema de segurança.
Você não silencia o oxímetro porque a especificidade é ruim. Você silencia porque o VPP é ruim. Qualquer pessoa já sabe intuitivamente qual das duas métricas governa o comportamento humano — só não sabe que ela tem nome. Falta exigir que os artigos reportem justamente a que todo mundo já usa na prática.
Escolher a Métrica É um Ato Clínico
O mérito real do estudo de Ha e colaboradores não é a rede neural. É que eles reportaram Average Precision como métrica primária num desfecho de 3% de prevalência, e reportaram o par sensibilidade/VPP em vez de sensibilidade/especificidade. Isso é raro, e é a razão pela qual dá para confiar na leitura deles.
Mas a escolha da métrica não é um detalhe de reporte. Ela codifica uma decisão sobre o que o modelo é. Otimizar sensibilidade produz um instrumento de rastreio. Otimizar VPP produz um instrumento de decisão. São produtos clínicos diferentes, com fluxos diferentes e responsabilidades diferentes — e a literatura trata a escolha como se fosse preferência estatística.
A consequência de arquitetura que quase ninguém tira
Sensibilidade 88% com VPP 32% não é um classificador mal ajustado. É a assinatura de qualquer modelo otimizado sobre desfecho raro e caro de perder. Você não conserta isso mexendo no limiar — mexer no limiar apenas troca uma métrica pela outra ao longo da mesma curva.
A saída é estrutural: cascata de dois estágios. Estágio 1 barato e sensível, que varre todo mundo e aceita falso-positivo. Estágio 2 caro e específico, acionado apenas sobre os positivos do estágio 1 — e o estágio 2 pode perfeitamente ser humano. Uma reavaliação estruturada de enfermagem em cinco minutos é um classificador de alta especificidade que já existe, já é validado e já está no serviço.
Um exemplo concreto: veredito contra convocação
Existem duas formas de o mesmo modelo, com exatamente a mesma performance estatística, aparecer na tela da triagem.
Como veredito: “Risco alto de necessidade de terapia intensiva.” Para isso ser aceito, o VPP precisa ser alto — senão a equipe descobre em três semanas que a máquina erra duas em cada três e para de olhar. Com VPP de 32%, esse produto morre.
Como convocação: “Reavaliação estruturada sugerida em 15 minutos.” Para isso ser aceito, o VPP não precisa ser alto — precisa apenas que o custo da reavaliação seja baixo. Cinco minutos de enfermagem, 82 vezes a cada mil atendimentos. Com VPP de 32%, esse produto funciona.
Mesmo modelo. Mesmos números. Um morre, o outro vive. A diferença não está no algoritmo — está em qual ato clínico ele solicita.
⚠ Onde a cascata ainda pode falhar
A cascata só funciona se o estágio 2 for de fato mais específico que o estágio 1 e tiver custo marginal baixo. Se a reavaliação estruturada não estiver protocolada, o que a convocação produz é 82 interrupções por mil atendimentos sem ganho de informação — o mesmo alarm fatigue, com etiqueta diferente. O estágio 2 é o produto; o modelo é só o gatilho.
A Camada Regulatória — CFM 2.454
A Resolução CFM nº 2.454/2026 classifica sistemas de IA por nível de risco — baixo, médio, alto ou inaceitável — considerando impacto em direitos fundamentais, complexidade do modelo, grau de autonomia e sensibilidade dos dados. O período de adaptação de 180 dias encerra em 26 de agosto de 2026, e o dever de classificar recai sobre a instituição que implementa, não sobre o fornecedor.
O detalhe que fecha o argumento deste artigo
Releia os critérios e note o segundo deles: grau de autonomia. Um sistema que emite veredito de prioridade sobre criança em fila de pronto-socorro tem autonomia decisória alta e impacto direto sobre acesso ao cuidado — dificilmente escapa de risco alto, com todo o ônus que isso carrega: validação documentada, supervisão médica sobre a saída e registro em prontuário do apoio da IA à decisão.
Um sistema que apenas convoca reavaliação humana tem autonomia decisória substancialmente menor: ele não decide nada, ele agenda um olhar. A classificação plausível cai para médio, possivelmente baixo. O mesmo modelo, com a mesma performance, muda de classe regulatória conforme o ato clínico que ele solicita. Isso não é uma brecha — é a norma funcionando como pretendido: ela regula autonomia, não acurácia.
⚠ O que está em uso hoje e ninguém classificou
A resolução alcança ferramentas já em operação. Isso inclui a categoria que nenhum serviço está mapeando: o LLM de uso geral acessado informalmente por membros da equipe dentro do fluxo assistencial. Em 26 de agosto, isso é um sistema de IA não classificado dentro de uma instituição médica. A resposta defensável não é proibir — é inventariar e classificar antes do prazo.
A Linha do Tempo
Da métrica publicada à classe de risco na sua instituição
Por que a decisão que define o destino de um modelo preditivo pediátrico é tomada antes da primeira linha de código — e depois da última.
O Que Fazer na Segunda-Feira
Da leitura crítica de um artigo à classificação do seu próprio sistema
Aprenda a Julgar um Modelo Antes de Confiar Nele
A metodologia AIMED forma médicos que constroem — não apenas consomem — ferramentas de IA clínica. Escolha de métrica sob desbalanceamento, arquitetura em cascata, desenho do ato clínico e classificação de risco CFM 2.454 fazem parte do currículo, porque são a mesma decisão vista de ângulos diferentes.
Considerações Finais
O modelo de Ha e colaboradores é bom. Average Precision de 0,60 sobre uma linha de base de 0,03 é vinte vezes o acaso, em 886 mil atendimentos, usando apenas dados de triagem. Não há nada a desmerecer no trabalho — pelo contrário, ele é exemplar justamente por reportar a métrica difícil quando poderia ter reportado a fácil.
O que há a desmerecer é o hábito de leitura que domina o meio médico brasileiro: aceitar a AUROC como veredito de qualidade em desfechos que são, quase todos, raros. A pergunta que importa não é “esse modelo tem AUC alta?”, e sim “qual é a prevalência, e o que acontece com a equipe quando esse alerta disparar pela terceira vez na mesma noite?”.
💡 Connecting the Dots: o ativo de autoridade aqui não é a Average Precision de 0,60 — é o fato de que a AUROC é matematicamente insensível à prevalência, e portanto o mesmo 0,95 descreve um modelo com VPP de 90% e outro com VPP de 10%. Todo mundo cita a AUROC; quase ninguém no meio médico brasileiro traduz que ela não sabe se o desfecho é raro, e por isso não sabe nada sobre o que vai acontecer no plantão. Mas o segundo salto é o que quase ninguém dá: se sensibilidade alta com VPP baixo é a assinatura inevitável do desfecho raro, então a variável de projeto não é o algoritmo — é o ato clínico que a saída solicita. “Risco alto” exige VPP que o desfecho raro não permite entregar; “reavaliação em 15 minutos” não exige. E como a CFM 2.454 classifica por grau de autonomia e não por acurácia, reescrever aquela frase na tela reduz simultaneamente a exigência estatística e a classe de risco regulatória. É engenharia clínica e engenharia regulatória sendo a mesma decisão — e essa decisão não se aprende em curso de machine learning, porque exige saber o que acontece com a enfermagem no terceiro mês, nem em residência, porque exige saber por que a curva precisão-recall existe. É exatamente essa interseção, e não o acesso ao modelo, que constitui o fosso técnico de quem constrói IA clínica no Brasil.
Referências
- Ha T, Kappy B, Chamberlain JM, McKinley KW. Early Prediction of Critical Care Interventions From Pediatric Emergency Department Triage. Hosp Pediatr. 2026 Ago;16(8):e618–e625. (886.183 visitas; 26.721 desfechos, 3,0%; rede neural AP 0,60 IC95% 0,59–0,61; sensibilidade 88%, VPP 32%; contrafactual na Tabela 4) Disponível em: https://doi.org/10.1542/hpeds.2025-009127
- Davis J, Goadrich M. The Relationship Between Precision-Recall and ROC Curves. Proceedings of the 23rd International Conference on Machine Learning (ICML). 2006:233–240. Disponível em: https://doi.org/10.1145/1143844.1143874
- Saito T, Rehmsmeier M. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLoS One. 2015;10(3):e0118432. (acesso aberto; ver p. 5 para a linha de base da PRC, y = P/(P+N), e Fig 2B para a demonstração visual) Disponível em: https://doi.org/10.1371/journal.pone.0118432
- Conselho Federal de Medicina. Resolução CFM nº 2.454, de 11 de fevereiro de 2026 — Normatiza o uso da inteligência artificial na medicina. DOU 2026 fev 27; Ed. 39, Seção 1. Disponível em: https://sistemas.cfm.org.br/normas/arquivos/resolucoes/BR/2026/2454_2026.pdf
- Jiang X, Yang N, Shi T, et al. Deciphering the “non-verbal” code: A preliminary exploration of multimodal large language models for neonatal pain recognition. Digit Health. 2026;12:20552076261473729. (mesma assimetria sensibilidade/especificidade em domínio distinto) Disponível em: https://doi.org/10.1177/20552076261473729
- Lonsdale H, Patel K, Domenico H, et al. Development and external validation of the NEO-READY model to predict date of discharge among premature neonatal intensive care patients. J Perinatol. 2026. Disponível em: https://doi.org/10.1038/s41372-026-02827-2
- Artificial Intelligence in Pediatric Cardiac Intensive Care: Clinical Applications, Implementation Challenges, and Future Directions. Curr Treat Options Pediatr. 2026. Disponível em: https://link.springer.com/article/10.1007/s40746-026-00374-8
The AI That Finds the Critically Ill Child in the ED Queue — and the Metric That Makes Every Model Look Better Than It Is
A model trained on 886,000 pediatric encounters identifies, using triage data alone, 88% of the children who will need a critical care intervention — and would triple the share of intermediate-risk patients evaluated in time. That is headline material. But there is a fact almost no clinical article translates: the metric that made these models famous is blind to what matters most in triage — how rare the event is. Understanding why AUROC lies for rare outcomes, and what to put in its place, is what separates someone who consumes a predictive score from someone who can judge one.
📅 Published July 27, 2026
Table of Contents
From the emergency queue to the regulatory risk class — why choosing the metric is choosing the clinical outcome
- 1. Why This Matters for Those on the Front Line
- 2. What the Study Delivered
- 3. The Numbers, Unfiltered
- 4. The Trap: The Metric That Cannot Feel Rarity
- 5. Choosing the Metric Is a Clinical Act
- 6. The Regulatory Layer — CFM 2.454
- 7. The Timeline
- 8. What to Do on Monday
- 9. Final Considerations
- 10. References
For readers outside healthcare — the essentials in 60 seconds
This article works on two layers. The technical one, for people who work in emergency departments or build models; and the general one, which requires only curiosity. Every core idea appears twice: once with a hospital example, once with an everyday one. If the first does not land, skip to the second — you will not miss anything.
Triage is what happens when you arrive at an emergency department: someone decides who is seen first. Not order of arrival, but order of severity. Brazil uses the Manchester Protocol, which sorts patients into five colors — red is now, yellow is soon, green can wait. The United States uses the ESI, which does the same with numbers 1 to 5.
The respiratory support acronyms ahead are simply degrees of help with breathing, from lightest to heaviest: HFNC (high-flow oxygen through the nose) and continuous nebulization are the light ones; CPAP and BiPAP are masks that push air under pressure; intubation is the tube in the windpipe, with a machine breathing for the patient. A vasopressor is the medication that holds up blood pressure when the body can no longer do it alone.
Prevalence is the fraction of people in whom something happens. “A prevalence of 3%” means that, of every 100 encounters, 3 end in that outcome. It is the most important word in this article — keep hold of it.
Why This Matters for Those on the Front Line
Every pediatric triage system has a purgatory. In the American system it is ESI level 3; in the Manchester system, which is ours in Brazil, it is yellow. It is the band where the child is not obvious enough to be red and not well enough to be green. It is where the bronchiolitis that will tire out in four hours lives, alongside the dehydration that has not yet closed the capillary refill, and the early sepsis that still has only a fever and nothing else.
The problem with tiered risk classification is not that it errs — it is that it saturates. A five-level system has, by construction, five degrees of resolution; and most of the movement happens inside a single level. The triage nurse knows this. The attending knows this. And that is exactly the gap a predictive model proposes to fill: not to replace Manchester, but to order what is stacked inside the yellow.
A study published on July 20, 2026 did precisely that, at scale. And the technical answer it gave is good. But what makes this paper worth a critical read is not the model — it is the metric the authors chose to report it with, and why that choice is itself a clinical decision.
What the Study Delivered
A retrospective cohort from 2016 to 2024, at a large urban academic pediatric emergency department in the United States with roughly 85,000 annual encounters. 886,183 encounters. The outcome occurred in 26,721 visits, or 3.0% of the total.
What exactly counts as the outcome — read this before comparing to your unit
The authors defined “need for critical care interventions” as the use of life-saving medications or respiratory support within 48 hours of ED presentation. Medications: continuous infusion of vasopressors, insulin, or terbutaline. Respiratory support: intubation, BiPAP, CPAP, heliox, high-flow nasal cannula (HFNC), and at least 2 hours of continuous nebulized albuterol. It also counts what was done on an inpatient unit and what was done at a return ED visit within 48 hours.
That is broader than it sounds. HFNC and two hours of continuous albuterol happen routinely in the observation area and on the ward — they are not, to a pediatric intensivist, “critical care.” A good share of that 3% is asthma and bronchiolitis receiving standard treatment, not a child on the verge of collapse. When you read the number, adjust the expectation: the model predicts the need to escalate support, not imminent arrest.
Here the authors did the right thing and deserve credit: they ran a more stringent secondary outcome, restricted to intubation, BiPAP, CPAP, and heliox — excluding HFNC and continuous albuterol — and performance held. They also tested a shorter 8-hour window, with similar results. Those are two sensitivity analyses most papers skip.
The authors trained six algorithms using exclusively information available at the moment of triage. No laboratory results, no evolution, no reassessment — only what the triage nurse holds in the first minutes. The neural network performed best.
The number that matters is not the model’s
The authors ran a counterfactual simulation: what would happen to time-to-physician if the score were used alongside the ESI rather than in place of it? In the ESI 3 band — the purgatory — the proportion of critical care patients evaluated in time would jump from 23.3% to 75.0%, and median time-to-pediatrician would fall from 34 to 10 minutes. For ESI 2, from 48.7% to 87.1% (19 → 12 min). For ESI 4, from 10.3% to 60.7% (62 → 7 min).
Note the design: the model reclassifies no one. It reorders within the class. That is an architectural choice that preserves the validated instrument the team already uses and adds resolution exactly where that instrument is blind. It is worth more as a lesson in clinical engineering than the AUC of any of the six algorithms.
A concrete example: what “predicting” means here
It is 9 p.m. on a winter Saturday. Reception holds 40 children waiting, 26 of them classified yellow. Manchester has already done its job: it separated the 3 reds and the 11 greens. What remains are 26 yellows, mutually indistinguishable on the screen, ordered by arrival time.
The model changes no one’s color. It reorders the 26 — and puts at the top the 5-month-old infant with a respiratory rate in the high percentile for age and a saturation of 93%, who arrived 40 minutes after a teenager with abdominal pain. None of that information is new. All of it was on the triage form. The model merely did the arithmetic nobody has time to do at 9 p.m. on a Saturday.
⚠ The counterfactual is the weakest link — and it is worth understanding why
The simulation rule is this: each patient who actually received a critical care intervention is assigned the earliest evaluation timestamp among all non-critical patients of the same ESI waiting concurrently. “Timely evaluation” means having been seen ahead of every non-critical patient at that same level.
Note what the rule does not model: the false positives. With a PPV of 32%, for every 1,000 encounters the model raises roughly 82 flags and only 26 are real — the other 56 are children who would also be pushed to the front of the queue, competing for exactly the same priority slots. The queue has finite capacity: there is only one “next available pediatrician.” The simulation grants the benefit to the true positives without charging them the competition from the false ones.
There is a productive irony here, and it is not in the paper’s limitations section: the same work that honestly reports a PPV of 32% runs a counterfactual that behaves, in practice, as if the PPV were far higher. This does not invalidate the study — but it turns the 23.3% → 75.0% into the ceiling of what the information could buy, not what implementation would deliver. Add to that what the authors themselves concede: single center, no prospective validation, with a declared need for recalibration at other sites.
⚠ The transfer blocker to Brazil that almost no one will notice
It sits in a single sentence of the limitations section: “our models depend on natural language processing of nursing narratives.” The model does not run on structured vital signs — it reads the free text the nurse writes at triage.
That changes everything for us. First, it is NLP trained on English, over American nursing documentation conventions. Second, and more serious: Brazilian Manchester triage is heavily structured around discriminators, and the quality and volume of free text vary enormously between services — in many, it is one line. The variable carrying the most signal in the original model is precisely the one that barely exists in the Brazilian workflow. This is not a recalibration case: it is a case for retraining on a different feature base.
The Numbers, Unfiltered
Before the table, five concepts. They sound dry, but each answers a question you already ask in your head on call — just without the technical name. Worth reading even if statistics is not your terrain: this article’s entire argument fits here.
A note on the “95% CI” appearing in the table: it is the 95% confidence interval — the range within which the true value probably lies. A narrow range, like the study’s 0.59–0.61, means a precise estimate, a consequence of the sample being enormous.
The metric decoder (start with the threshold — everything depends on it)
Threshold — “from what point do I call?” A model does not return “yes” or “no.” It returns a continuous number, something like 0.17 or 0.64. Someone has to decide from which value that becomes an alert on the screen. That cut-off point is the threshold.
In the hospital: the saturation at which you decide to call the attending is a threshold. Raise it from 92% to 94% and you call earlier and more often: you catch nearly every child who was going to worsen, and you call a lot of people for nothing. Drop it to 88% and you call less often and are almost always right — but you miss the ones who were just starting to sink.
Outside it: it is the strictness setting on your spam filter. Too tight, and real email lands in the trash. Too loose, and junk floods your inbox. You cannot have both — and no adjustment of the filter fixes it, because the problem is not the filter, it is having to pick one point on a single scale.
Sensitivity and positive predictive value are exactly that balance, and moving the threshold trades one for the other. There is no setting that improves both at once: it is the same rope, pulled from opposite ends. Hold on to this — it is the piece that carries this article’s conclusion.
Sensitivity — “of those who would need the ICU, how many does the model catch?” Sensitivity of 88% means that, of 100 children who will receive a critical care intervention, the model raises the alert on 88. The other 12 slip through. It is the metric that hurts to miss.
Positive predictive value (PPV) — “when it calls me, how often is it real?” A PPV of 32% means that, of every 100 alerts fired, 32 are children who really will need care. The other 68 are not. That is the number that decides whether the team is still looking at the alert in week three.
AUROC — “can it tell which of two children is sicker?” Draw two children at random, one who will deteriorate and one who will not. The AUROC is the probability the model gives the higher score to the one who deteriorates. It is a test of pairwise ranking. And here is the key this entire article pursues: that test does not change whether the event is common or vanishingly rare. AUROC is mathematically insensitive to prevalence.
Average Precision (AP) — “of the times it called me, how many were worth it?” It is the mean PPV, computed by sweeping every possible threshold at once. Instead of giving you the PPV at one specific cut-off, it gives you the behavior of the whole model. And, unlike AUROC, AP feels how rare the event is.
The “baseline” of a metric is the score a useless model earns — the one flipping a coin. It works as a ruler: without knowing the useless model’s score, you cannot tell whether the good model’s score is good. And here is the difference that decides everything, demonstrated by Saito and Rehmsmeier:
The AUROC baseline is always 0.50, whether the outcome happens in half of patients or in one per thousand. The AP baseline is the prevalence itself — the authors show it equals exactly P/(P+N), the fraction of positive cases in the total. In plain shift-floor terms: for an outcome occurring in 3% of children, flipping a coin earns an AP of 0.03. Not 0.50 — 0.03.
That is why the study’s AP of 0.60 means twenty times better than chance. Had it been an AUROC of 0.60, it would be barely above a coin flip. Same number, opposite readings — because the rulers are different.
| Metric | Reported value | Baseline | Source | Reading |
|---|---|---|---|---|
| Cohort size | 886,183 encounters (2016–2024) | — | Ha et al. · Hosp Pediatr · 2026 | ↑ robust |
| Outcome prevalence | 26,721 visits · 3.0% | — | Ha et al. · Hosp Pediatr · 2026 | ⚠ rare event |
| Average Precision (neural network) | 0.60 (95% CI 0.59–0.61) | 0.03 | Ha et al. · Hosp Pediatr · 2026 | ↑ 20× chance |
| Sensitivity | 88% (87–89%) | — | Ha et al. · Hosp Pediatr · 2026 | ↑ good |
| Positive predictive value | 32% (31–32%) | 3% | Ha et al. · Hosp Pediatr · 2026 | ↓ 2 in 3 alerts are false |
| Specificity (derived, not reported) | ≈ 94% | — | Own calculation from sens. and PPV | → see box below |
| ESI 3 evaluated in time | 23.3% → 75.0% | — | Ha et al. · counterfactual analysis | ↑ ceiling, not floor |
The arithmetic that reorganizes the whole table
Specificity was not reported in the study abstract, but it can be derived. With a prevalence of 3%, sensitivity of 88%, and PPV of 32%, for every 1,000 encounters: 30 children will need an intervention; the model catches 26 of them; to reach a PPV of 32% it must fire roughly 82 alerts in total; therefore 56 are false positives. That leaves 914 non-events correctly kept silent out of 970. Specificity ≈ 94%.
Now read the following two sentences, which describe exactly the same classifier: “specificity of 94%” and “two out of every three alerts are false.” The first passes any committee. The second describes the nursing staff’s actual experience in month three. There is no contradiction — there are two different questions. Specificity asks what happens to those who are fine; PPV asks what happens to those who were called. For rare events the denominators are brutally different in size, which is why the two answers diverge so sharply.
If you have ever been through an airport, you have lived this. The metal detector catches essentially every weapon that passes through it — very high sensitivity, exactly as it should be. It also stays silent for the overwhelming majority of passengers — very high specificity. And still, when it beeps, it is almost never a weapon: it is a belt buckle, a key, a coin forgotten in a pocket. All three things are true at once, and they only seem contradictory until you notice that weapons are vanishingly rare among passengers. That is exactly the structure of this article’s problem — and the reason nobody switches off the airport detector while everybody switches off the hospital alert: at the airport, the cost of a false alarm is ten seconds with a wand. In the hospital, when the cost of a false alarm is high, the alert dies.
The Trap: The Metric That Cannot Feel Rarity
Here is the fact almost no clinical popularizer translates: AUROC does not change when prevalence changes. This is neither a practical limitation nor a sampling bias — it is a mathematical property. AUROC is built on sensitivity and specificity, and both are computed within each group: sensitivity looks only at the sick, specificity only at the well. Neither knows the ratio between the groups.
The consequence is uncomfortable. The same AUROC value describes a model with 90% PPV and a model with 10% PPV, depending solely on how rare the outcome is. If you read “AUROC 0.95” without knowing the prevalence, you know absolutely nothing about how the alert will behave on shift.
Why this is epidemic in pediatric literature
Because nearly every serious pediatric outcome is rare. Cardiac arrest, sepsis, need for ventilation, death — all with single-digit prevalence. That is precisely the regime in which AUROC is most generous and least informative. A mediocre model on a 2%-prevalence outcome easily produces an AUROC of 0.90, because it is easy to separate the overwhelming majority of well children from the few sick ones; the hard part is getting which of the few right.
This critique is neither mine nor new. Davis and Goadrich formally demonstrated, in 2006, that the precision-recall curve is more informative on imbalanced datasets — jargon for “almost everyone is negative and very few are positive,” which is the exact description of any serious pediatric outcome — and that optimizing the area under the ROC does not guarantee optimizing the area under the precision-recall curve. Saito and Rehmsmeier, in 2015, put it plainly in their own abstract: the visual interpretability of ROC plots on imbalanced datasets “can be deceptive with respect to conclusions about the reliability of classification performance, owing to an intuitive but wrong interpretation of specificity.” Twenty years of established methodology, nearly 4,700 citations, and pediatric clinical literature still reports AUROC for 3%-prevalence outcomes.
⚠ Where my own argument has limits: prevalence is not the same as spectrum
I need to be precise here, because an attentive reader will find this on their own. Saying “AUROC is insensitive to prevalence” is a statement about the construction of the metric: sensitivity and specificity are ratios computed within each group, which is why the ROC baseline does not move when the proportion between groups changes.
That does not mean the measured AUROC will be the same in any service. If the new population has sicker cases, milder cases, or a different mix, the scores the model hands out spread differently — and the AUROC you measure in practice shifts with them. That is the spectrum effect: the proportion of sick patients did not change, the kind of sick patient did. These are two distinct problems that often appear together and get conflated: prevalence invariance is a mathematical property of the metric; the spectrum effect is a change in the population.
The study under discussion illustrates the second problem well: the cohort is 61.8% non-Hispanic Black children and 23.2% Hispanic or Latino, with 74% covered by public insurance, at a single hospital in Washington, DC. It is a real and well-described population — and it is another population.
The practical consequence, however, is the same in both cases, and it reinforces rather than weakens the argument: a published AUROC does not let you anticipate how many false alerts your team will receive. In the first case because the metric does not carry that information; in the second because the population is not the same. Either way, the number you need is the PPV computed at your prevalence.
⚠ The risk is not the isolated false positive — it is the death of the alert
Two false alerts out of three are not a statistical inconvenience. They are the mechanism of alarm fatigue, and alarm fatigue has documented clinical consequences: an alert nobody looks at protects less than no alert at all, because it consumes attention and produces a false sense of coverage. A predictive triage system with low PPV and the wrong presentation does not degrade gracefully — it is mentally switched off by the team within weeks, and stays switched on on the screen for months.
A concrete example: the saturation alarm you already silenced today
You know this metric without knowing its name. The pulse oximeter on bed 4 beeps every twenty minutes. The overwhelming majority of the time it is a displaced sensor, a child who moved a hand, an artifact. That alarm’s specificity is very high — it stays quiet most of the time, relative to all the minutes the child is fine. Its positive predictive value is very low — when it screams, it is almost never a real desaturation.
Outside the hospital, the same: think of the car alarm that goes off in the street at three in the morning. Nobody gets up. Nobody looks out the window. Not because the alarm is never right — but because it is almost always wrong, and the human brain learns that within weeks. A car alarm has excellent sensitivity for break-ins and negligible positive predictive value, which is why it became urban noise instead of a security system.
You do not silence the oximeter because specificity is poor. You silence it because PPV is poor. Anyone already knows intuitively which of the two metrics governs human behavior — they just do not know it has a name. What is missing is demanding that papers report the very one everyone already uses in practice.
Choosing the Metric Is a Clinical Act
The real merit of Ha and colleagues’ study is not the neural network. It is that they reported Average Precision as the primary metric for a 3%-prevalence outcome, and reported the sensitivity/PPV pair instead of sensitivity/specificity. That is rare, and it is why their reading can be trusted.
But the choice of metric is not a reporting detail. It encodes a decision about what the model is. Optimizing sensitivity produces a screening instrument. Optimizing PPV produces a decision instrument. These are different clinical products, with different workflows and different responsibilities — and the literature treats the choice as if it were statistical preference.
The architectural consequence almost nobody draws
Sensitivity of 88% with a PPV of 32% is not a badly tuned classifier. It is the signature of any model optimized on an outcome that is rare and costly to miss. You do not fix that by moving the threshold — moving the threshold merely trades one metric for the other along the same curve.
The way out is structural: a two-stage cascade. Stage 1 cheap and sensitive, sweeping everyone and accepting false positives. Stage 2 expensive and specific, triggered only on stage 1’s positives — and stage 2 can perfectly well be human. A five-minute structured nursing reassessment is a high-specificity classifier that already exists, is already validated, and is already in the service.
A concrete example: verdict versus summons
There are two ways the same model, with exactly the same statistical performance, can appear on the triage screen.
As a verdict: “High risk of needing critical care.” For that to be accepted, PPV must be high — otherwise the team discovers within three weeks that the machine is wrong two times out of three and stops looking. With a PPV of 32%, that product dies.
As a summons: “Structured reassessment suggested within 15 minutes.” For that to be accepted, PPV need not be high — it only requires that the cost of reassessment be low. Five minutes of nursing time, 82 times per thousand encounters. With a PPV of 32%, that product works.
Same model. Same numbers. One dies, the other lives. The difference is not in the algorithm — it is in which clinical act the output requests.
⚠ Where the cascade can still fail
The cascade only works if stage 2 is genuinely more specific than stage 1 and carries low marginal cost. If the structured reassessment is not protocolized, what the summons produces is 82 interruptions per thousand encounters with no information gain — the same alarm fatigue, under a different label. Stage 2 is the product; the model is only the trigger.
The Regulatory Layer — CFM 2.454
CFM Resolution No. 2,454/2026 classifies AI systems by risk level — low, medium, high, or unacceptable — considering impact on fundamental rights, model complexity, degree of autonomy, and data sensitivity. The 180-day adaptation period ends on August 26, 2026, and the duty to classify falls on the implementing institution, not the vendor.
The detail that closes this article’s argument
Reread the criteria and note the second one: degree of autonomy. A system issuing a priority verdict about a child in an emergency queue has high decisional autonomy and direct impact on access to care — it hardly escapes high risk, with all the burden that carries: documented validation, physician oversight of the output, and a chart entry recording AI support for the decision.
A system that merely summons a human reassessment has substantially lower decisional autonomy: it decides nothing, it schedules a look. The plausible classification drops to medium, possibly low. The same model, with the same performance, changes regulatory class according to the clinical act it requests. This is not a loophole — it is the norm working as intended: it regulates autonomy, not accuracy.
⚠ What is in use today and nobody has classified
The resolution reaches tools already in operation. That includes the category no service is mapping: the general-purpose LLM accessed informally by team members within the care workflow. On August 26, that is an unclassified AI system inside a medical institution. The defensible response is not to ban it — it is to inventory and classify it before the deadline.
The Timeline
From the published metric to the risk class in your institution
Why the decision that determines a pediatric predictive model’s fate is made before the first line of code — and after the last.
What to Do on Monday
From critically reading a paper to classifying your own system
Learn to Judge a Model Before Trusting It
The AIMED methodology trains physicians who build — not merely consume — clinical AI tools. Metric selection under class imbalance, cascade architecture, clinical act design, and CFM 2.454 risk classification are part of the curriculum, because they are the same decision seen from different angles.
Final Considerations
Ha and colleagues’ model is good. An Average Precision of 0.60 against a baseline of 0.03 is twenty times chance, across 886,000 encounters, using triage data alone. There is nothing to diminish in the work — on the contrary, it is exemplary precisely because it reported the hard metric when it could have reported the easy one.
What there is to diminish is the reading habit that dominates clinical medicine: accepting AUROC as a verdict of quality for outcomes that are, almost all of them, rare. The question that matters is not “does this model have a high AUC?” but “what is the prevalence, and what happens to the team when that alert fires for the third time in the same night?”.
💡 Connecting the Dots: the authority asset here is not the Average Precision of 0.60 — it is the fact that AUROC is mathematically insensitive to prevalence, and therefore the same 0.95 describes a model with 90% PPV and another with 10% PPV. Everyone cites the AUROC; almost no one in Brazilian medicine translates that it does not know whether the outcome is rare, and therefore knows nothing about what will happen on shift. But the second leap is the one almost nobody makes: if high sensitivity with low PPV is the inevitable signature of a rare outcome, then the design variable is not the algorithm — it is the clinical act the output requests. “High risk” demands a PPV that a rare outcome cannot deliver; “reassessment within 15 minutes” does not. And because CFM 2.454 classifies by degree of autonomy rather than accuracy, rewriting that sentence on the screen simultaneously lowers the statistical requirement and the regulatory risk class. Clinical engineering and regulatory engineering turn out to be the same decision — and that decision is not taught in a machine learning course, because it requires knowing what happens to the nursing staff in month three, nor in residency, because it requires knowing why the precision-recall curve exists. It is exactly that intersection, and not access to the model, that constitutes the technical moat for anyone building clinical AI in Brazil.
References
- Ha T, Kappy B, Chamberlain JM, McKinley KW. Early Prediction of Critical Care Interventions From Pediatric Emergency Department Triage. Hosp Pediatr. 2026 Aug;16(8):e618–e625. (886,183 visits; 26,721 outcomes, 3.0%; neural network AP 0.60, 95% CI 0.59–0.61; sensitivity 88%, PPV 32%; counterfactual in Table 4) Available at: https://doi.org/10.1542/hpeds.2025-009127
- Davis J, Goadrich M. The Relationship Between Precision-Recall and ROC Curves. Proceedings of the 23rd International Conference on Machine Learning (ICML). 2006:233–240. Available at: https://doi.org/10.1145/1143844.1143874
- Saito T, Rehmsmeier M. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLoS One. 2015;10(3):e0118432. (open access; see p. 5 for the PRC baseline, y = P/(P+N), and Fig 2B for the visual demonstration) Available at: https://doi.org/10.1371/journal.pone.0118432
- Federal Council of Medicine (CFM). Resolution CFM No. 2,454, of February 11, 2026 — Regulating the use of artificial intelligence in medicine. Official Gazette 2026 Feb 27; Ed. 39, Sec. 1. Available at: https://sistemas.cfm.org.br/normas/arquivos/resolucoes/BR/2026/2454_2026.pdf
- Jiang X, Yang N, Shi T, et al. Deciphering the “non-verbal” code: A preliminary exploration of multimodal large language models for neonatal pain recognition. Digit Health. 2026;12:20552076261473729. (same sensitivity/specificity asymmetry in a distinct domain) Available at: https://doi.org/10.1177/20552076261473729
- Lonsdale H, Patel K, Domenico H, et al. Development and external validation of the NEO-READY model to predict date of discharge among premature neonatal intensive care patients. J Perinatol. 2026. Available at: https://doi.org/10.1038/s41372-026-02827-2
- Artificial Intelligence in Pediatric Cardiac Intensive Care: Clinical Applications, Implementation Challenges, and Future Directions. Curr Treat Options Pediatr. 2026. Available at: https://link.springer.com/article/10.1007/s40746-026-00374-8
