{"id":2805,"date":"2026-07-28T18:09:27","date_gmt":"2026-07-28T21:09:27","guid":{"rendered":"https:\/\/inovamed.pro\/?p=2805"},"modified":"2026-07-28T18:36:16","modified_gmt":"2026-07-28T21:36:16","slug":"metrica-auroc-triagem-pediatrica-average-precision","status":"publish","type":"post","link":"https:\/\/inovamed.pro\/?p=2805","title":{"rendered":"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9"},"content":{"rendered":"<div id=\"cfm-root\" class=\"ino-fullbleed\" data-deadline=\"2026-08-26\">\n<div class=\"nav-buttons\">\n<a href=\"#top-pt\" id=\"nav-btn-top\" class=\"nav-btn\" title=\"Topo\">\u2191<\/a><br \/>\n<a href=\"#indice-pt\" id=\"nav-btn-toc\" class=\"nav-btn secondary\" title=\"\u00cdndice\">\ud83e\udded<\/a>\n<\/div>\n<div class=\"language-toggle\">\n<a href=\"#\" class=\"lang-btn active\" id=\"btn-pt\">\ud83c\udde7\ud83c\uddf7 PT<\/a><br \/>\n<a href=\"#\" class=\"lang-btn\" id=\"btn-en\">\ud83c\uddfa\ud83c\uddf8 EN<\/a>\n<\/div>\n<p><!-- ===================== CONTE\u00daDO PT ===================== --><\/p>\n<div id=\"content-pt\">\n<section class=\"hero\" id=\"top-pt\">\n<span class=\"article-tag\">Emerg\u00eancia Pedi\u00e1trica \u00b7 Metodologia \u00b7 Machine Learning<\/span><\/p>\n<h1>A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9<\/h1>\n<p class=\"subtitle\">Um modelo treinado em 886 mil atendimentos pedi\u00e1tricos identifica, s\u00f3 com dados da triagem, 88% das crian\u00e7as que v\u00e3o precisar de interven\u00e7\u00e3o de terapia intensiva \u2014 e triplicaria a fra\u00e7\u00e3o de pacientes de risco intermedi\u00e1rio avaliados a tempo. \u00c9 o tipo de n\u00famero que rende manchete. Mas h\u00e1 um dado que quase nenhum artigo brasileiro traduz: <strong>a m\u00e9trica que consagrou esses modelos \u00e9 cega para o que mais importa na triagem \u2014 a raridade do evento<\/strong>. Entender por que a AUROC mente em desfecho raro, e o que colocar no lugar dela, \u00e9 o que separa quem consome escore preditivo de quem sabe julg\u00e1-lo.<\/p>\n<p class=\"meta-info\">Dr. Mbula Luzingu Barros | M\u00e9dico Pediatra Intensivista \u00b7 25 anos de UTI Pedi\u00e1trica | Consultor em IA na Sa\u00fade | Fundador da INOVAMED | Criador da Metodologia AIMED<\/p>\n<p class=\"date\">\ud83d\udcc5 Publicado em 27 de julho de 2026<\/p>\n<div class=\"deadline-badge\">\u23f1 Vig\u00eancia plena da CFM 2.454 em: <span class=\"countdown-days\">&#8212;<\/span> dias \u2014 26 de agosto de 2026<\/div>\n<\/section>\n<div class=\"container\">\n<div class=\"toc\" id=\"indice-pt\">\n<h2>Navegue pelo Artigo<\/h2>\n<p class=\"toc-subtitle\">Da fila do pronto-socorro \u00e0 classe de risco regulat\u00f3ria \u2014 por que escolher a m\u00e9trica \u00e9 escolher o desfecho cl\u00ednico<\/p>\n<ul>\n<li><a href=\"#s1-pt\">1. Por Que Isso Importa Para Quem Trabalha na Ponta<\/a><\/li>\n<li><a href=\"#s2-pt\">2. O Que o Estudo Entregou<\/a><\/li>\n<li><a href=\"#s3-pt\">3. Os N\u00fameros, Sem Filtro<\/a><\/li>\n<li><a href=\"#s4-pt\">4. A Armadilha: A M\u00e9trica Que N\u00e3o Sente a Raridade<\/a><\/li>\n<li><a href=\"#s5-pt\">5. Escolher a M\u00e9trica \u00c9 um Ato Cl\u00ednico<\/a><\/li>\n<li><a href=\"#s6-pt\">6. A Camada Regulat\u00f3ria \u2014 CFM 2.454<\/a><\/li>\n<li><a href=\"#s7-pt\">7. A Linha do Tempo<\/a><\/li>\n<li><a href=\"#s8-pt\">8. O Que Fazer na Segunda-Feira<\/a><\/li>\n<li><a href=\"#final-pt\">9. Considera\u00e7\u00f5es Finais<\/a><\/li>\n<li><a href=\"#ref-pt\">10. Refer\u00eancias<\/a><\/li>\n<\/ul>\n<\/div>\n<div class=\"science-box\">\n<h4>Para quem n\u00e3o \u00e9 da \u00e1rea \u2014 o essencial em 60 segundos<\/h4>\n<p>Este artigo funciona em duas camadas. A t\u00e9cnica, para quem trabalha em pronto-socorro ou constr\u00f3i modelos; e a geral, que s\u00f3 exige curiosidade. Toda ideia central aparece duas vezes: uma com exemplo de hospital, outra com exemplo do dia a dia. Se a primeira n\u00e3o fizer sentido, pule para a segunda \u2014 n\u00e3o vai perder nada.<\/p>\n<p><strong>Triagem<\/strong> \u00e9 o que acontece quando voc\u00ea chega a um pronto-socorro: algu\u00e9m decide quem \u00e9 atendido primeiro. N\u00e3o \u00e9 ordem de chegada, \u00e9 ordem de gravidade. No Brasil usamos o <strong>Protocolo de Manchester<\/strong>, que separa os pacientes em cinco cores \u2014 vermelho \u00e9 agora, amarelo \u00e9 logo, verde pode esperar. Nos Estados Unidos usam o <strong>ESI<\/strong>, que faz o mesmo com n\u00fameros de 1 a 5.<\/p>\n<p><strong>As siglas de suporte respirat\u00f3rio<\/strong> que v\u00e3o aparecer s\u00e3o apenas graus de ajuda para respirar, do mais leve ao mais pesado: <em>HFNC<\/em> (oxig\u00eanio em alto fluxo pelo nariz) e nebuliza\u00e7\u00e3o cont\u00ednua s\u00e3o os leves; <em>CPAP<\/em> e <em>BiPAP<\/em> s\u00e3o m\u00e1scaras que empurram ar com press\u00e3o; <em>intuba\u00e7\u00e3o<\/em> \u00e9 o tubo na traqueia, com o aparelho respirando pelo paciente. <em>Vasopressor<\/em> \u00e9 a medica\u00e7\u00e3o que sustenta a press\u00e3o arterial quando o corpo n\u00e3o consegue mais.<\/p>\n<p><strong>Preval\u00eancia<\/strong> \u00e9 a fra\u00e7\u00e3o de pessoas em quem algo acontece. &#8220;Preval\u00eancia de 3%&#8221; significa que, de cada 100 atendimentos, 3 terminam naquele desfecho. \u00c9 a palavra mais importante do artigo \u2014 guarde essa.<\/p>\n<\/div>\n<section class=\"section\" id=\"s1-pt\">\n<h2 class=\"section-title\">Por Que Isso Importa Para Quem Trabalha na Ponta<\/h2>\n<div class=\"intro-box\">\n<p>Toda triagem pedi\u00e1trica tem um purgat\u00f3rio. No sistema americano \u00e9 o <strong>ESI n\u00edvel 3<\/strong>; no Manchester, que \u00e9 o nosso, \u00e9 o <strong>amarelo<\/strong>. \u00c9 a faixa onde a crian\u00e7a n\u00e3o est\u00e1 \u00f3bvia o suficiente para virar vermelho e n\u00e3o est\u00e1 bem o suficiente para virar verde. \u00c9 onde mora a bronquiolite que vai cansar em quatro horas, a desidrata\u00e7\u00e3o que ainda n\u00e3o fechou o tempo de enchimento capilar, a sepse precoce que ainda tem febre e nada mais.<\/p>\n<p>O problema da classifica\u00e7\u00e3o de risco por n\u00edveis n\u00e3o \u00e9 que ela erre \u2014 \u00e9 que ela <strong>satura<\/strong>. Um sistema de cinco n\u00edveis tem, por constru\u00e7\u00e3o, cinco graus de resolu\u00e7\u00e3o; e a maior parte do movimento acontece dentro de um \u00fanico n\u00edvel. O enfermeiro classificador sabe disso. O plantonista sabe disso. E \u00e9 exatamente essa a lacuna que um modelo preditivo se prop\u00f5e a preencher: n\u00e3o substituir o Manchester, mas <strong>ordenar o que est\u00e1 empilhado dentro do amarelo<\/strong>.<\/p>\n<p>Um estudo publicado em 20 de julho de 2026 fez precisamente isso, em escala. E a resposta t\u00e9cnica que ele deu \u00e9 boa. Mas o que torna esse artigo digno de leitura cr\u00edtica n\u00e3o \u00e9 o modelo \u2014 \u00e9 a <em>m\u00e9trica<\/em> que os autores escolheram para report\u00e1-lo, e o motivo de essa escolha ser, ela pr\u00f3pria, uma decis\u00e3o cl\u00ednica.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s2-pt\">\n<h2 class=\"section-title\">O Que o Estudo Entregou<\/h2>\n<div class=\"detailed-section\">\n<p>Coorte retrospectiva de 2016 a 2024, em um pronto-socorro pedi\u00e1trico acad\u00eamico urbano de grande porte nos Estados Unidos, com cerca de 85 mil atendimentos anuais. <strong>886.183 atendimentos.<\/strong> O desfecho ocorreu em <strong>26.721 visitas, ou 3,0% do total<\/strong>.<\/p>\n<h4>O que exatamente conta como desfecho \u2014 leia antes de comparar com a sua unidade<\/h4>\n<p>Os autores definiram &#8220;necessidade de interven\u00e7\u00e3o de suporte avan\u00e7ado&#8221; como <strong>uso de medica\u00e7\u00e3o de suporte \u00e0 vida ou suporte respirat\u00f3rio dentro de 48 horas da chegada ao pronto-socorro<\/strong>. Medica\u00e7\u00e3o: infus\u00e3o cont\u00ednua de vasopressor, insulina ou terbutalina. Suporte respirat\u00f3rio: intuba\u00e7\u00e3o, BiPAP, CPAP, heliox, <strong>cateter nasal de alto fluxo (HFNC)<\/strong> e <strong>pelo menos 2 horas de nebuliza\u00e7\u00e3o cont\u00ednua de salbutamol<\/strong>. Conta tamb\u00e9m o que foi feito na enfermaria e o que foi feito em retorno ao pronto-socorro dentro de 48 horas.<\/p>\n<p>Isso \u00e9 mais largo do que soa. HFNC e salbutamol cont\u00ednuo por duas horas, no Brasil, acontecem rotineiramente na sala de observa\u00e7\u00e3o e na enfermaria \u2014 n\u00e3o s\u00e3o, para n\u00f3s, &#8220;terapia intensiva&#8221;. <strong>Boa parte dos 3% \u00e9 asma e bronquiolite recebendo tratamento padr\u00e3o, n\u00e3o crian\u00e7a \u00e0 beira do colapso.<\/strong> Ao ler o n\u00famero, ajuste a expectativa: o modelo prev\u00ea necessidade de escalonamento de suporte, n\u00e3o imin\u00eancia de parada.<\/p>\n<p>Aqui os autores fizeram a coisa certa e merecem cr\u00e9dito: rodaram um <strong>desfecho secund\u00e1rio mais estrito<\/strong>, s\u00f3 com intuba\u00e7\u00e3o, BiPAP, CPAP e heliox \u2014 excluindo HFNC e salbutamol cont\u00ednuo \u2014 e o desempenho se manteve. Tamb\u00e9m testaram uma janela mais curta, de 8 horas, com resultado semelhante. S\u00e3o duas an\u00e1lises de sensibilidade que a maioria dos artigos n\u00e3o faz.<\/p>\n<p>Os autores treinaram seis algoritmos usando <strong>exclusivamente informa\u00e7\u00e3o dispon\u00edvel no momento da triagem<\/strong>. Nada de exame laboratorial, nada de evolu\u00e7\u00e3o, nada de reavalia\u00e7\u00e3o \u2014 s\u00f3 o que o enfermeiro classificador tem em m\u00e3os nos primeiros minutos. A rede neural teve o melhor desempenho.<\/p>\n<h4>O n\u00famero que interessa n\u00e3o \u00e9 o do modelo<\/h4>\n<p>Os autores fizeram uma simula\u00e7\u00e3o contrafactual: o que aconteceria com o tempo at\u00e9 a avalia\u00e7\u00e3o m\u00e9dica se o escore fosse usado <em>junto<\/em> com o ESI, e n\u00e3o no lugar dele? Na faixa ESI 3 \u2014 o purgat\u00f3rio \u2014, a propor\u00e7\u00e3o de pacientes de suporte avan\u00e7ado avaliados em tempo h\u00e1bil saltaria de <strong>23,3% para 75,0%<\/strong>, e a mediana de tempo at\u00e9 o pediatra cairia de <strong>34 para 10 minutos<\/strong>. No ESI 2, de 48,7% para 87,1% (19 \u2192 12 min). No ESI 4, de 10,3% para 60,7% (62 \u2192 7 min).<\/p>\n<p>Repare no desenho: o modelo n\u00e3o reclassifica ningu\u00e9m. Ele <em>reordena dentro da classe<\/em>. \u00c9 uma escolha de arquitetura que preserva o instrumento validado que a equipe j\u00e1 usa e adiciona resolu\u00e7\u00e3o onde o instrumento \u00e9 cego. Vale mais como li\u00e7\u00e3o de engenharia cl\u00ednica do que a AUC de qualquer um dos seis algoritmos.<\/p>\n<\/div>\n<div class=\"highlight-box\">\n<h4>Um exemplo concreto: o que &#8220;prever&#8221; significa aqui<\/h4>\n<p>S\u00e3o 21h de um s\u00e1bado de inverno. A recep\u00e7\u00e3o tem 40 crian\u00e7as esperando, 26 delas classificadas como amarelo. O Manchester j\u00e1 fez o trabalho dele: separou os 3 vermelhos e os 11 verdes. Sobram 26 amarelos indistingu\u00edveis entre si na tela, ordenados por ordem de chegada.<\/p>\n<p>O modelo n\u00e3o muda a cor de ningu\u00e9m. Ele reordena os 26 \u2014 e coloca no topo a lactente de 5 meses com frequ\u00eancia respirat\u00f3ria no percentil alto para a idade e satura\u00e7\u00e3o de 93%, que chegou 40 minutos depois de um adolescente com dor abdominal. <strong>Nenhuma dessas informa\u00e7\u00f5es \u00e9 nova.<\/strong> Todas estavam na ficha de triagem. O modelo s\u00f3 fez a aritm\u00e9tica que ningu\u00e9m tem tempo de fazer \u00e0s 21h de s\u00e1bado.<\/p>\n<\/div>\n<div class=\"warning-box\">\n<h4>\u26a0 O contrafactual \u00e9 o elo mais fraco \u2014 e vale entender por qu\u00ea<\/h4>\n<p>A regra da simula\u00e7\u00e3o \u00e9 esta: a cada paciente que <em>de fato<\/em> recebeu suporte avan\u00e7ado, atribui-se o hor\u00e1rio de atendimento mais cedo entre todos os pacientes <em>n\u00e3o cr\u00edticos<\/em> do mesmo ESI que esperavam simultaneamente. &#8220;Avaliado em tempo h\u00e1bil&#8221; significa ter sido visto antes de todos os n\u00e3o cr\u00edticos daquele mesmo n\u00edvel.<\/p>\n<p>Repare no que a regra <strong>n\u00e3o<\/strong> modela: os <em>falsos-positivos<\/em>. Com VPP de 32%, a cada 1.000 atendimentos o modelo levanta cerca de 82 sinaliza\u00e7\u00f5es e apenas 26 s\u00e3o reais \u2014 as outras 56 s\u00e3o crian\u00e7as que tamb\u00e9m seriam empurradas para a frente da fila, disputando exatamente as mesmas vagas de prioridade. A fila tem capacidade finita: s\u00f3 existe um &#8220;pr\u00f3ximo pediatra dispon\u00edvel&#8221;. A simula\u00e7\u00e3o concede o benef\u00edcio aos verdadeiros-positivos sem cobrar deles a competi\u00e7\u00e3o dos falsos.<\/p>\n<p>H\u00e1 uma ironia produtiva aqui, e ela n\u00e3o est\u00e1 na se\u00e7\u00e3o de limita\u00e7\u00f5es do artigo: <strong>o mesmo trabalho que reporta honestamente um VPP de 32% roda um contrafactual que se comporta, na pr\u00e1tica, como se o VPP fosse muito maior.<\/strong> Isso n\u00e3o invalida o estudo \u2014 mas transforma o 23,3% \u2192 75,0% no <em>teto<\/em> do que a informa\u00e7\u00e3o poderia comprar, n\u00e3o no que a implementa\u00e7\u00e3o entregaria. Some a isso o que os pr\u00f3prios autores admitem: centro \u00fanico, sem valida\u00e7\u00e3o prospectiva, com necessidade declarada de recalibra\u00e7\u00e3o para outros servi\u00e7os.<\/p>\n<\/div>\n<div class=\"alert-box\">\n<h4>\u26a0 O bloqueio de transfer\u00eancia para o Brasil que quase ningu\u00e9m vai notar<\/h4>\n<p>Est\u00e1 numa \u00fanica frase da se\u00e7\u00e3o de limita\u00e7\u00f5es: <em>&#8220;nossos modelos dependem de processamento de linguagem natural das narrativas de enfermagem&#8221;<\/em>. O modelo n\u00e3o roda sobre sinais vitais estruturados \u2014 ele l\u00ea o <strong>texto livre que o enfermeiro escreve na triagem<\/strong>.<\/p>\n<p>Isso muda tudo para n\u00f3s. Primeiro, \u00e9 NLP treinado em ingl\u00eas, sobre conven\u00e7\u00f5es de registro de enfermagem americanas. Segundo, e mais grave: a classifica\u00e7\u00e3o de risco brasileira pelo Manchester \u00e9 fortemente estruturada em discriminadores, e a qualidade e o volume do texto livre variam enormemente entre servi\u00e7os \u2014 em muitos, \u00e9 uma linha. <strong>A vari\u00e1vel que mais carrega sinal no modelo original \u00e9 justamente a que menos existe no fluxo brasileiro.<\/strong> N\u00e3o \u00e9 caso de recalibrar: \u00e9 caso de retreinar sobre outra base de features.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s3-pt\">\n<h2 class=\"section-title\">Os N\u00fameros, Sem Filtro<\/h2>\n<p>Antes da tabela, cinco conceitos. Eles parecem \u00e1ridos, mas cada um responde a uma pergunta que voc\u00ea j\u00e1 faz de cabe\u00e7a no plant\u00e3o \u2014 s\u00f3 que sem o nome t\u00e9cnico. Vale ler mesmo se estat\u00edstica n\u00e3o \u00e9 o seu terreno: o argumento inteiro deste artigo cabe aqui.<\/p>\n<p><em>Uma nota sobre o &#8220;IC95%&#8221; que aparece na tabela: \u00e9 o intervalo de confian\u00e7a de 95% \u2014 a faixa dentro da qual o valor verdadeiro provavelmente est\u00e1. Faixa estreita, como o 0,59\u20130,61 do estudo, significa estimativa precisa, efeito de a amostra ser enorme.<\/em><\/p>\n<div class=\"science-box\">\n<h4>O tradutor de m\u00e9tricas (comece pelo limiar \u2014 tudo depende dele)<\/h4>\n<p><strong>Limiar \u2014 &#8220;a partir de que ponto eu chamo?&#8221;<\/strong> Um modelo n\u00e3o devolve &#8220;sim&#8221; ou &#8220;n\u00e3o&#8221;. Devolve um n\u00famero cont\u00ednuo, tipo 0,17 ou 0,64. Algu\u00e9m precisa decidir a partir de qual valor aquilo vira um alerta na tela. Esse ponto de corte \u00e9 o <strong>limiar<\/strong>.<\/p>\n<p><em>No hospital:<\/em> a satura\u00e7\u00e3o em que voc\u00ea decide chamar o plantonista \u00e9 um limiar. Se voc\u00ea sobe de 92% para 94%, chama mais cedo e mais vezes: pega quase todas as crian\u00e7as que iam piorar, e chama muita gente \u00e0 toa. Se desce para 88%, chama menos e quase sempre com raz\u00e3o \u2014 mas deixa passar as que estavam come\u00e7ando a afundar.<\/p>\n<p><em>Fora dele:<\/em> \u00e9 o ajuste de rigor do seu filtro de spam. Apertado demais, e-mail de verdade cai na lixeira. Frouxo demais, propaganda entope a caixa de entrada. Voc\u00ea n\u00e3o consegue os dois \u2014 e nenhum ajuste do filtro resolve, porque o problema n\u00e3o est\u00e1 no filtro, est\u00e1 em ter que escolher um ponto na mesma escala.<\/p>\n<p><strong>Sensibilidade e valor preditivo positivo s\u00e3o exatamente esse balan\u00e7o, e mexer no limiar troca um pelo outro.<\/strong> N\u00e3o existe ajuste que melhore os dois ao mesmo tempo: \u00e9 a mesma corda, puxada de pontas opostas. Guarde isso \u2014 \u00e9 a pe\u00e7a que sustenta a conclus\u00e3o deste artigo.<\/p>\n<p><strong>Sensibilidade \u2014 &#8220;de quem ia precisar de UTI, quantos o modelo pega?&#8221;<\/strong> Sensibilidade 88% quer dizer que, de 100 crian\u00e7as que v\u00e3o receber interven\u00e7\u00e3o de terapia intensiva, o modelo acende o alerta em 88. As outras 12 passam. \u00c9 a m\u00e9trica que d\u00f3i errar.<\/p>\n<p><strong>Valor preditivo positivo (VPP) \u2014 &#8220;quando ele me chama, quantas vezes \u00e9 de verdade?&#8221;<\/strong> VPP 32% quer dizer que, de cada 100 alertas disparados, 32 s\u00e3o crian\u00e7as que realmente v\u00e3o precisar. As outras 68 n\u00e3o. Esse \u00e9 o n\u00famero que decide se a equipe vai continuar olhando para o alerta depois da terceira semana.<\/p>\n<p><strong>AUROC \u2014 &#8220;ele sabe dizer qual das duas est\u00e1 mais grave?&#8221;<\/strong> Sorteie duas crian\u00e7as, uma que vai complicar e outra que n\u00e3o. A AUROC \u00e9 a probabilidade de o modelo dar a nota mais alta para a que complica. \u00c9 uma prova de <em>ordena\u00e7\u00e3o entre pares<\/em>. E aqui est\u00e1 a chave que este artigo inteiro persegue: <strong>essa prova n\u00e3o muda se o evento for comum ou rar\u00edssimo<\/strong>. A AUROC \u00e9 matematicamente insens\u00edvel \u00e0 preval\u00eancia.<\/p>\n<p><strong>Average Precision (AP) \u2014 &#8220;das vezes que ele me chamou, quantas valeram a pena?&#8221;<\/strong> \u00c9 o VPP m\u00e9dio, calculado percorrendo todos os limiares poss\u00edveis de uma vez s\u00f3. Em vez de te dar o VPP de um ponto de corte espec\u00edfico, te d\u00e1 o comportamento do modelo inteiro. E, diferente da AUROC, a AP <em>sente<\/em> a raridade do evento.<\/p>\n<p><strong>A &#8220;linha de base&#8221; de uma m\u00e9trica<\/strong> \u00e9 a nota que um modelo <em>in\u00fatil<\/em> tira \u2014 aquele que sorteia no cara ou coroa. Serve de r\u00e9gua: sem saber a nota do in\u00fatil, voc\u00ea n\u00e3o sabe se a nota do bom \u00e9 boa. E aqui est\u00e1 a diferen\u00e7a que decide tudo, demonstrada por Saito e Rehmsmeier:<\/p>\n<p>A <strong>linha de base da AUROC \u00e9 sempre 0,50<\/strong>, n\u00e3o importa se o desfecho acontece em metade dos pacientes ou em um a cada mil. A <strong>linha de base da AP \u00e9 a pr\u00f3pria preval\u00eancia<\/strong> \u2014 os autores demonstram que ela vale exatamente P\/(P+N), a fra\u00e7\u00e3o de casos positivos no total. Em portugu\u00eas de plant\u00e3o: num desfecho que ocorre em 3% das crian\u00e7as, chutar no cara ou coroa rende AP de <strong>0,03<\/strong>. N\u00e3o 0,50 \u2014 0,03.<\/p>\n<p>\u00c9 por isso que a AP de 0,60 do estudo significa <strong>vinte vezes melhor que o acaso<\/strong>. Se fosse AUROC de 0,60, seria pouco acima de cara ou coroa. Mesmo n\u00famero, leituras opostas \u2014 porque as r\u00e9guas s\u00e3o diferentes.<\/p>\n<\/div>\n<div class=\"table-wrapper\">\n<table class=\"data-table\">\n<thead>\n<tr>\n<th>M\u00e9trica<\/th>\n<th>Valor reportado<\/th>\n<th>Linha de base<\/th>\n<th>Fonte<\/th>\n<th>Leitura<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Volume da coorte<\/td>\n<td>886.183 atendimentos (2016\u20132024)<\/td>\n<td>\u2014<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-baixo\">\u2191 robusto<\/span><\/td>\n<\/tr>\n<tr>\n<td>Preval\u00eancia do desfecho<\/td>\n<td>26.721 visitas \u00b7 <strong>3,0%<\/strong><\/td>\n<td>\u2014<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-alto\">\u26a0 evento raro<\/span><\/td>\n<\/tr>\n<tr>\n<td><strong>Average Precision<\/strong> (rede neural)<\/td>\n<td><strong>0,60<\/strong> (IC95% 0,59\u20130,61)<\/td>\n<td>0,03<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-baixo\">\u2191 20\u00d7 o acaso<\/span><\/td>\n<\/tr>\n<tr>\n<td>Sensibilidade<\/td>\n<td>88% (87\u201389%)<\/td>\n<td>\u2014<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-medio\">\u2191 bom<\/span><\/td>\n<\/tr>\n<tr>\n<td>Valor preditivo positivo<\/td>\n<td>32% (31\u201332%)<\/td>\n<td>3%<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-alto\">\u2193 2 em 3 alertas s\u00e3o falsos<\/span><\/td>\n<\/tr>\n<tr>\n<td>Especificidade <em>(derivada, n\u00e3o reportada)<\/em><\/td>\n<td>\u2248 94%<\/td>\n<td>\u2014<\/td>\n<td>C\u00e1lculo pr\u00f3prio a partir de sens. e VPP<\/td>\n<td><span class=\"status-badge s-inaceitavel\">\u2192 ver box abaixo<\/span><\/td>\n<\/tr>\n<tr>\n<td>ESI 3 avaliados em tempo h\u00e1bil<\/td>\n<td>23,3% \u2192 <strong>75,0%<\/strong><\/td>\n<td>\u2014<\/td>\n<td>Ha et al. \u00b7 an\u00e1lise contrafactual<\/td>\n<td><span class=\"status-badge s-medio\">\u2191 teto, n\u00e3o piso<\/span><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<div class=\"science-box\">\n<h4>A conta que reorganiza a tabela inteira<\/h4>\n<p>A especificidade n\u00e3o foi reportada no resumo do estudo, mas d\u00e1 para deriv\u00e1-la. Com preval\u00eancia de 3%, sensibilidade de 88% e VPP de 32%, a cada <strong>1.000 atendimentos<\/strong>: 30 crian\u00e7as v\u00e3o precisar de interven\u00e7\u00e3o; o modelo acerta 26 delas; para chegar a um VPP de 32% ele precisa disparar cerca de <strong>82 alertas no total<\/strong>; logo, 56 s\u00e3o falsos positivos. Sobram 914 n\u00e3o-eventos corretamente silenciados de um total de 970. <strong>Especificidade \u2248 94%.<\/strong><\/p>\n<p>Agora leia as duas frases seguintes, que descrevem exatamente o mesmo classificador: <em>&#8220;especificidade de 94%&#8221;<\/em> e <em>&#8220;dois de cada tr\u00eas alertas s\u00e3o falsos&#8221;<\/em>. A primeira passa em qualquer comit\u00ea. A segunda descreve a experi\u00eancia real da enfermagem no terceiro m\u00eas de uso. <strong>N\u00e3o h\u00e1 contradi\u00e7\u00e3o \u2014 h\u00e1 duas perguntas diferentes.<\/strong> Especificidade pergunta o que acontece com quem est\u00e1 bem; VPP pergunta o que acontece com quem foi chamado. Em evento raro, os denominadores s\u00e3o de tamanhos brutalmente diferentes, e por isso as duas respostas divergem tanto.<\/p>\n<p><strong>Se voc\u00ea j\u00e1 passou por um aeroporto, j\u00e1 viveu isso.<\/strong> O detector de metal pega praticamente toda arma que passa por ele \u2014 sensibilidade alt\u00edssima, e \u00e9 assim que tem que ser. Ele tamb\u00e9m fica calado para a esmagadora maioria dos passageiros \u2014 especificidade alt\u00edssima. E ainda assim, quando ele apita, quase nunca \u00e9 arma: \u00e9 fivela de cinto, chave, moeda esquecida no bolso. <strong>As tr\u00eas coisas s\u00e3o verdadeiras ao mesmo tempo<\/strong>, e s\u00f3 parecem contradit\u00f3rias at\u00e9 voc\u00ea notar que arma \u00e9 rar\u00edssima entre passageiros. \u00c9 exatamente a estrutura do problema deste artigo \u2014 e a raz\u00e3o pela qual ningu\u00e9m desliga o detector de aeroporto, mas todo mundo desliga o alerta do hospital: no aeroporto, o custo do falso alarme \u00e9 passar o bast\u00e3o em algu\u00e9m por dez segundos. No hospital, quando o custo do falso alarme \u00e9 alto, o alerta morre.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s4-pt\">\n<h2 class=\"section-title\">A Armadilha: A M\u00e9trica Que N\u00e3o Sente a Raridade<\/h2>\n<div class=\"detailed-section\">\n<p>Aqui est\u00e1 o dado que quase nenhum divulgador brasileiro traduz: <strong>a AUROC n\u00e3o muda quando a preval\u00eancia muda.<\/strong> Isso n\u00e3o \u00e9 uma limita\u00e7\u00e3o pr\u00e1tica nem um vi\u00e9s de amostragem \u2014 \u00e9 uma propriedade matem\u00e1tica. A AUROC \u00e9 constru\u00edda sobre sensibilidade e especificidade, e ambas s\u00e3o calculadas <em>dentro<\/em> de cada grupo: sensibilidade s\u00f3 olha os doentes, especificidade s\u00f3 olha os sadios. Nenhuma das duas sabe a propor\u00e7\u00e3o entre os grupos.<\/p>\n<p>A consequ\u00eancia \u00e9 desconfort\u00e1vel. <strong>O mesmo valor de AUROC descreve um modelo com VPP de 90% e um modelo com VPP de 10%<\/strong>, a depender apenas de qu\u00e3o raro \u00e9 o desfecho. Se voc\u00ea l\u00ea &#8220;AUROC 0,95&#8221; sem saber a preval\u00eancia, voc\u00ea n\u00e3o sabe absolutamente nada sobre como o alerta vai se comportar no plant\u00e3o.<\/p>\n<h4>Por que isso \u00e9 epid\u00eamico na literatura pedi\u00e1trica<\/h4>\n<p>Porque quase todo desfecho grave em pediatria \u00e9 raro. Parada cardiorrespirat\u00f3ria, sepse, necessidade de ventila\u00e7\u00e3o, \u00f3bito \u2014 todos com preval\u00eancia de um d\u00edgito. \u00c9 precisamente o regime em que a AUROC \u00e9 mais generosa e menos informativa. Um modelo med\u00edocre num desfecho de 2% de preval\u00eancia facilmente produz uma AUROC de 0,90, porque \u00e9 f\u00e1cil separar a maioria esmagadora de sadios dos poucos doentes; o dif\u00edcil \u00e9 acertar <em>qual<\/em> dos poucos.<\/p>\n<p>Essa cr\u00edtica n\u00e3o \u00e9 minha e nem \u00e9 nova. Davis e Goadrich demonstraram formalmente, em 2006, que a curva precis\u00e3o-recall \u00e9 mais informativa em bases <em>desbalanceadas<\/em> \u2014 jarg\u00e3o para &#8220;quase todo mundo \u00e9 negativo e pouqu\u00edssimos s\u00e3o positivos&#8221;, que \u00e9 a descri\u00e7\u00e3o exata de qualquer desfecho grave em pediatria \u2014, e que otimizar a \u00e1rea sob a ROC n\u00e3o garante otimizar a \u00e1rea sob a curva precis\u00e3o-recall. Saito e Rehmsmeier, em 2015, foram diretos ao ponto no pr\u00f3prio resumo: a interpretabilidade visual da ROC em dados desbalanceados <em>&#8220;pode ser enganosa quanto \u00e0s conclus\u00f5es sobre a confiabilidade do desempenho de classifica\u00e7\u00e3o, devido a uma interpreta\u00e7\u00e3o intuitiva por\u00e9m errada da especificidade&#8221;<\/em>. <strong>Vinte anos de metodologia estabelecida, quase 4.700 cita\u00e7\u00f5es, e a literatura cl\u00ednica pedi\u00e1trica ainda reporta AUROC em desfecho de 3%.<\/strong><\/p>\n<\/div>\n<div class=\"warning-box\">\n<h4>\u26a0 Onde meu pr\u00f3prio argumento tem limite: preval\u00eancia n\u00e3o \u00e9 a mesma coisa que espectro<\/h4>\n<p>Preciso ser preciso, porque um leitor atento vai encontrar isto sozinho. Dizer que &#8220;a AUROC \u00e9 insens\u00edvel \u00e0 preval\u00eancia&#8221; \u00e9 uma afirma\u00e7\u00e3o sobre a <strong>constru\u00e7\u00e3o da m\u00e9trica<\/strong>: sensibilidade e especificidade s\u00e3o raz\u00f5es calculadas dentro de cada grupo, e por isso a linha de base da ROC n\u00e3o se move quando a propor\u00e7\u00e3o entre os grupos muda.<\/p>\n<p>Isso <em>n\u00e3o<\/em> significa que a AUROC medida seja a mesma em qualquer servi\u00e7o. Se a nova popula\u00e7\u00e3o tiver casos mais graves, mais leves, ou uma mistura diferente, as notas que o modelo distribui se espalham de outro jeito \u2014 e a AUROC medida na pr\u00e1tica muda junto. \u00c9 o <strong>efeito de espectro<\/strong>: n\u00e3o mudou a propor\u00e7\u00e3o de doentes, mudou o <em>tipo<\/em> de doente. S\u00e3o dois problemas distintos que frequentemente aparecem juntos e s\u00e3o confundidos: <em>invari\u00e2ncia \u00e0 preval\u00eancia<\/em> \u00e9 propriedade matem\u00e1tica da m\u00e9trica; <em>efeito de espectro<\/em> \u00e9 mudan\u00e7a da popula\u00e7\u00e3o.<\/p>\n<p>O estudo que discutimos ilustra bem o segundo problema: a coorte \u00e9 61,8% de crian\u00e7as negras n\u00e3o hisp\u00e2nicas, 23,2% hisp\u00e2nicas ou latinas, com 74% cobertas por seguro p\u00fablico, num \u00fanico hospital de Washington. \u00c9 uma popula\u00e7\u00e3o real e bem descrita \u2014 e \u00e9 <em>outra<\/em> popula\u00e7\u00e3o.<\/p>\n<p>A consequ\u00eancia pr\u00e1tica, por\u00e9m, \u00e9 a mesma nos dois casos, e refor\u00e7a o argumento em vez de enfraquec\u00ea-lo: <strong>a AUROC publicada n\u00e3o permite antecipar quantos alertas falsos a sua equipe vai receber.<\/strong> No primeiro caso porque a m\u00e9trica n\u00e3o carrega essa informa\u00e7\u00e3o; no segundo porque a popula\u00e7\u00e3o n\u00e3o \u00e9 a mesma. Em qualquer dos dois, o n\u00famero que voc\u00ea precisa \u00e9 o VPP calculado na sua preval\u00eancia.<\/p>\n<\/div>\n<div class=\"alert-box\">\n<h4>\u26a0 O risco n\u00e3o \u00e9 o falso positivo isolado \u2014 \u00e9 a morte do alerta<\/h4>\n<p>Dois de cada tr\u00eas alertas falsos n\u00e3o s\u00e3o um inconveniente estat\u00edstico. S\u00e3o o mecanismo do <em>alarm fatigue<\/em>, e o <em>alarm fatigue<\/em> tem desfecho cl\u00ednico documentado: o alerta que ningu\u00e9m mais olha protege menos que nenhum alerta, porque consome aten\u00e7\u00e3o e produz falsa sensa\u00e7\u00e3o de cobertura. Um sistema de triagem preditiva com VPP baixo e apresenta\u00e7\u00e3o errada n\u00e3o degrada graciosamente \u2014 ele \u00e9 desligado mentalmente pela equipe em semanas, e continua ligado na tela por meses.<\/p>\n<\/div>\n<div class=\"highlight-box\">\n<h4>Um exemplo concreto: o alarme de satura\u00e7\u00e3o que voc\u00ea j\u00e1 silenciou hoje<\/h4>\n<p>Voc\u00ea conhece essa m\u00e9trica sem saber o nome dela. O ox\u00edmetro do leito 4 apita a cada vinte minutos. Na esmagadora maioria das vezes \u00e9 o sensor deslocado, \u00e9 a crian\u00e7a que mexeu a m\u00e3o, \u00e9 artefato. A <em>especificidade<\/em> daquele alarme \u00e9 alt\u00edssima \u2014 ele fica calado a maior parte do tempo, em rela\u00e7\u00e3o a todos os minutos em que a crian\u00e7a est\u00e1 bem. O <em>valor preditivo positivo<\/em> dele \u00e9 baix\u00edssimo \u2014 quando ele grita, quase nunca \u00e9 dessatura\u00e7\u00e3o real.<\/p>\n<p><em>Fora do hospital, o mesmo:<\/em> pense no alarme de carro que dispara na rua \u00e0s tr\u00eas da manh\u00e3. Ningu\u00e9m levanta. Ningu\u00e9m olha pela janela. N\u00e3o \u00e9 porque o alarme nunca acerta \u2014 \u00e9 porque quase sempre erra, e o c\u00e9rebro humano aprende isso em poucas semanas. Um alarme de carro tem sensibilidade excelente para arrombamento e valor preditivo positivo desprez\u00edvel, e \u00e9 por isso que ele virou ru\u00eddo urbano em vez de sistema de seguran\u00e7a.<\/p>\n<p>Voc\u00ea n\u00e3o silencia o ox\u00edmetro porque a especificidade \u00e9 ruim. Voc\u00ea silencia porque o VPP \u00e9 ruim. <strong>Qualquer pessoa j\u00e1 sabe intuitivamente qual das duas m\u00e9tricas governa o comportamento humano \u2014 s\u00f3 n\u00e3o sabe que ela tem nome.<\/strong> Falta exigir que os artigos reportem justamente a que todo mundo j\u00e1 usa na pr\u00e1tica.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s5-pt\">\n<h2 class=\"section-title\">Escolher a M\u00e9trica \u00c9 um Ato Cl\u00ednico<\/h2>\n<div class=\"detailed-section\">\n<p>O m\u00e9rito real do estudo de Ha e colaboradores n\u00e3o \u00e9 a rede neural. \u00c9 que eles reportaram <strong>Average Precision<\/strong> como m\u00e9trica prim\u00e1ria num desfecho de 3% de preval\u00eancia, e reportaram o par sensibilidade\/VPP em vez de sensibilidade\/especificidade. Isso \u00e9 raro, e \u00e9 a raz\u00e3o pela qual d\u00e1 para confiar na leitura deles.<\/p>\n<p>Mas a escolha da m\u00e9trica n\u00e3o \u00e9 um detalhe de reporte. Ela codifica uma decis\u00e3o sobre <strong>o que o modelo \u00e9<\/strong>. Otimizar sensibilidade produz um instrumento de rastreio. Otimizar VPP produz um instrumento de decis\u00e3o. S\u00e3o produtos cl\u00ednicos diferentes, com fluxos diferentes e responsabilidades diferentes \u2014 e a literatura trata a escolha como se fosse prefer\u00eancia estat\u00edstica.<\/p>\n<h4>A consequ\u00eancia de arquitetura que quase ningu\u00e9m tira<\/h4>\n<p>Sensibilidade 88% com VPP 32% n\u00e3o \u00e9 um classificador mal ajustado. \u00c9 a <strong>assinatura<\/strong> de qualquer modelo otimizado sobre desfecho raro e caro de perder. Voc\u00ea n\u00e3o conserta isso mexendo no limiar \u2014 mexer no limiar apenas troca uma m\u00e9trica pela outra ao longo da mesma curva.<\/p>\n<p>A sa\u00edda \u00e9 estrutural: <strong>cascata de dois est\u00e1gios<\/strong>. Est\u00e1gio 1 barato e sens\u00edvel, que varre todo mundo e aceita falso-positivo. Est\u00e1gio 2 caro e espec\u00edfico, acionado <em>apenas<\/em> sobre os positivos do est\u00e1gio 1 \u2014 e o est\u00e1gio 2 pode perfeitamente ser humano. Uma reavalia\u00e7\u00e3o estruturada de enfermagem em cinco minutos \u00e9 um classificador de alta especificidade que j\u00e1 existe, j\u00e1 \u00e9 validado e j\u00e1 est\u00e1 no servi\u00e7o.<\/p>\n<\/div>\n<div class=\"highlight-box\">\n<h4>Um exemplo concreto: veredito contra convoca\u00e7\u00e3o<\/h4>\n<p>Existem duas formas de o mesmo modelo, com exatamente a mesma performance estat\u00edstica, aparecer na tela da triagem.<\/p>\n<p><strong>Como veredito:<\/strong> <em>&#8220;Risco alto de necessidade de terapia intensiva.&#8221;<\/em> Para isso ser aceito, o VPP precisa ser alto \u2014 sen\u00e3o a equipe descobre em tr\u00eas semanas que a m\u00e1quina erra duas em cada tr\u00eas e para de olhar. Com VPP de 32%, esse produto morre.<\/p>\n<p><strong>Como convoca\u00e7\u00e3o:<\/strong> <em>&#8220;Reavalia\u00e7\u00e3o estruturada sugerida em 15 minutos.&#8221;<\/em> Para isso ser aceito, o VPP <em>n\u00e3o precisa<\/em> ser alto \u2014 precisa apenas que o custo da reavalia\u00e7\u00e3o seja baixo. Cinco minutos de enfermagem, 82 vezes a cada mil atendimentos. Com VPP de 32%, esse produto funciona.<\/p>\n<p><strong>Mesmo modelo. Mesmos n\u00fameros. Um morre, o outro vive.<\/strong> A diferen\u00e7a n\u00e3o est\u00e1 no algoritmo \u2014 est\u00e1 em qual ato cl\u00ednico ele solicita.<\/p>\n<\/div>\n<div class=\"warning-box\">\n<h4>\u26a0 Onde a cascata ainda pode falhar<\/h4>\n<p>A cascata s\u00f3 funciona se o est\u00e1gio 2 for de fato mais espec\u00edfico que o est\u00e1gio 1 <em>e<\/em> tiver custo marginal baixo. Se a reavalia\u00e7\u00e3o estruturada n\u00e3o estiver protocolada, o que a convoca\u00e7\u00e3o produz \u00e9 82 interrup\u00e7\u00f5es por mil atendimentos sem ganho de informa\u00e7\u00e3o \u2014 o mesmo <em>alarm fatigue<\/em>, com etiqueta diferente. O est\u00e1gio 2 \u00e9 o produto; o modelo \u00e9 s\u00f3 o gatilho.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s6-pt\">\n<h2 class=\"section-title\">A Camada Regulat\u00f3ria \u2014 CFM 2.454<\/h2>\n<div class=\"detailed-section\">\n<p>A Resolu\u00e7\u00e3o CFM n\u00ba 2.454\/2026 classifica sistemas de IA por n\u00edvel de risco \u2014 baixo, m\u00e9dio, alto ou inaceit\u00e1vel \u2014 considerando impacto em direitos fundamentais, complexidade do modelo, <strong>grau de autonomia<\/strong> e sensibilidade dos dados. O per\u00edodo de adapta\u00e7\u00e3o de 180 dias encerra em <strong>26 de agosto de 2026<\/strong>, e o dever de classificar recai sobre a <em>institui\u00e7\u00e3o que implementa<\/em>, n\u00e3o sobre o fornecedor.<\/p>\n<h4>O detalhe que fecha o argumento deste artigo<\/h4>\n<p>Releia os crit\u00e9rios e note o segundo deles: <strong>grau de autonomia<\/strong>. Um sistema que emite veredito de prioridade sobre crian\u00e7a em fila de pronto-socorro tem autonomia decis\u00f3ria alta e impacto direto sobre acesso ao cuidado \u2014 dificilmente escapa de risco alto, com todo o \u00f4nus que isso carrega: valida\u00e7\u00e3o documentada, supervis\u00e3o m\u00e9dica sobre a sa\u00edda e registro em prontu\u00e1rio do apoio da IA \u00e0 decis\u00e3o.<\/p>\n<p>Um sistema que apenas <em>convoca reavalia\u00e7\u00e3o humana<\/em> tem autonomia decis\u00f3ria substancialmente menor: ele n\u00e3o decide nada, ele agenda um olhar. A classifica\u00e7\u00e3o plaus\u00edvel cai para m\u00e9dio, possivelmente baixo. <strong>O mesmo modelo, com a mesma performance, muda de classe regulat\u00f3ria conforme o ato cl\u00ednico que ele solicita.<\/strong> Isso n\u00e3o \u00e9 uma brecha \u2014 \u00e9 a norma funcionando como pretendido: ela regula autonomia, n\u00e3o acur\u00e1cia.<\/p>\n<\/div>\n<div class=\"alert-box\">\n<h4>\u26a0 O que est\u00e1 em uso hoje e ningu\u00e9m classificou<\/h4>\n<p>A resolu\u00e7\u00e3o alcan\u00e7a <strong>ferramentas j\u00e1 em opera\u00e7\u00e3o<\/strong>. Isso inclui a categoria que nenhum servi\u00e7o est\u00e1 mapeando: o LLM de uso geral acessado informalmente por membros da equipe dentro do fluxo assistencial. Em 26 de agosto, isso \u00e9 um sistema de IA n\u00e3o classificado dentro de uma institui\u00e7\u00e3o m\u00e9dica. A resposta defens\u00e1vel n\u00e3o \u00e9 proibir \u2014 \u00e9 inventariar e classificar antes do prazo.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s7-pt\">\n<h2 class=\"section-title\">A Linha do Tempo<\/h2>\n<div class=\"flow-box\">\n<h4>Da m\u00e9trica publicada \u00e0 classe de risco na sua institui\u00e7\u00e3o<\/h4>\n<p class=\"flow-intro\">Por que a decis\u00e3o que define o destino de um modelo preditivo pedi\u00e1trico \u00e9 tomada antes da primeira linha de c\u00f3digo \u2014 e depois da \u00faltima.<\/p>\n<div class=\"flow-steps\">\n<div class=\"flow-step\">\n<div class=\"step-label\">Publica\u00e7\u00e3o<\/div>\n<div class=\"step-title\">AUROC alta em evento raro<\/div>\n<div class=\"step-desc\">A m\u00e9trica insens\u00edvel \u00e0 preval\u00eancia produz a manchete. O VPP n\u00e3o aparece no resumo.<\/div>\n<\/div>\n<div class=\"flow-arrow\">\u279c<\/div>\n<div class=\"flow-step\">\n<div class=\"step-label\">Realidade<\/div>\n<div class=\"step-title\">Dois em cada tr\u00eas alertas falsos<\/div>\n<div class=\"step-desc\">Especificidade 94% e VPP 32% convivem. A equipe silencia o alerta em semanas.<\/div>\n<\/div>\n<div class=\"flow-arrow\">\u279c<\/div>\n<div class=\"flow-step current\">\n<div class=\"step-label\">Hoje \u00b7 Voc\u00ea est\u00e1 aqui<\/div>\n<div class=\"step-title\">Reformular o ato cl\u00ednico<\/div>\n<div class=\"step-desc\">Convoca\u00e7\u00e3o de reavalia\u00e7\u00e3o no lugar de veredito de prioridade. Cascata de dois est\u00e1gios.<\/div>\n<\/div>\n<div class=\"flow-arrow\">\u279c<\/div>\n<div class=\"flow-step\">\n<div class=\"step-label\">26\/08\/2026<\/div>\n<div class=\"step-title\">Classifica\u00e7\u00e3o CFM 2.454<\/div>\n<div class=\"step-desc\">Autonomia menor, classe de risco menor, \u00f4nus de conformidade menor. Dever da institui\u00e7\u00e3o.<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s8-pt\">\n<h2 class=\"section-title\">O Que Fazer na Segunda-Feira<\/h2>\n<div class=\"highlight-box\">\n<h4>Da leitura cr\u00edtica de um artigo \u00e0 classifica\u00e7\u00e3o do seu pr\u00f3prio sistema<\/h4>\n<div class=\"checklist-section\">\n<div class=\"check-subhead\">Ao ler qualquer artigo de modelo preditivo pedi\u00e1trico<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Procure a preval\u00eancia do desfecho antes de olhar a AUROC.<\/strong> Abaixo de 10%, trate AUROC isolada como n\u00e3o informativa e exija AUPRC ou Average Precision.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Confira se a AP vem acompanhada da linha de base.<\/strong> AP sem a preval\u00eancia ao lado \u00e9 um n\u00famero sem escala \u2014 a linha de base da AP <em>\u00e9<\/em> a preval\u00eancia.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Se s\u00f3 houver sensibilidade e especificidade, derive o VPP voc\u00ea mesmo.<\/strong> Precisa apenas da preval\u00eancia. \u00c9 a conta que revela quantos alertas falsos a equipe vai receber.<\/div>\n<\/div>\n<div class=\"check-subhead\">Antes de colocar qualquer modelo em produ\u00e7\u00e3o<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Decida se o produto \u00e9 rastreio ou decis\u00e3o<\/strong> \u2014 e reporte a m\u00e9trica correspondente. Rastreio responde por sensibilidade; decis\u00e3o responde por VPP.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Desenhe o est\u00e1gio 2 antes do est\u00e1gio 1.<\/strong> Se a reavalia\u00e7\u00e3o estruturada n\u00e3o estiver protocolada e cronometrada, o modelo s\u00f3 produz interrup\u00e7\u00e3o.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Escreva o texto exato que aparece na tela.<\/strong> &#8220;Risco alto&#8221; e &#8220;reavalia\u00e7\u00e3o sugerida em 15 minutos&#8221; t\u00eam exig\u00eancias de VPP e classes de risco regulat\u00f3rias diferentes.<\/div>\n<\/div>\n<div class=\"check-subhead\">Antes de 26 de agosto<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Inventarie os sistemas de IA em uso na sua unidade<\/strong> em quatro categorias: embarcados em equipamento, software diagn\u00f3stico contratado, LLM de uso geral acessado pela equipe, e modelos desenvolvidos internamente.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Classifique cada um pelo grau de autonomia<\/strong>, n\u00e3o pela acur\u00e1cia. \u00c9 esse o crit\u00e9rio da norma.<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/section>\n<div class=\"cta-section\">\n<h2>Aprenda a Julgar um Modelo Antes de Confiar Nele<\/h2>\n<p>A metodologia AIMED forma m\u00e9dicos que constroem \u2014 n\u00e3o apenas consomem \u2014 ferramentas de IA cl\u00ednica. Escolha de m\u00e9trica sob desbalanceamento, arquitetura em cascata, desenho do ato cl\u00ednico e classifica\u00e7\u00e3o de risco CFM 2.454 fazem parte do curr\u00edculo, porque s\u00e3o a mesma decis\u00e3o vista de \u00e2ngulos diferentes.<\/p>\n<p><a href=\"https:\/\/inovamed.pro\/?page_id=96\" class=\"cta-button\">Conhe\u00e7a o AIMED \u2192<\/a>\n<\/div>\n<section class=\"section\" id=\"final-pt\">\n<h2 class=\"section-title\">Considera\u00e7\u00f5es Finais<\/h2>\n<div class=\"intro-box\">\n<p>O modelo de Ha e colaboradores \u00e9 bom. Average Precision de 0,60 sobre uma linha de base de 0,03 \u00e9 vinte vezes o acaso, em 886 mil atendimentos, usando apenas dados de triagem. N\u00e3o h\u00e1 nada a desmerecer no trabalho \u2014 pelo contr\u00e1rio, ele \u00e9 exemplar justamente por reportar a m\u00e9trica dif\u00edcil quando poderia ter reportado a f\u00e1cil.<\/p>\n<p>O que h\u00e1 a desmerecer \u00e9 o h\u00e1bito de leitura que domina o meio m\u00e9dico brasileiro: aceitar a AUROC como veredito de qualidade em desfechos que s\u00e3o, quase todos, raros. A pergunta que importa n\u00e3o \u00e9 &#8220;esse modelo tem AUC alta?&#8221;, e sim <em>&#8220;qual \u00e9 a preval\u00eancia, e o que acontece com a equipe quando esse alerta disparar pela terceira vez na mesma noite?&#8221;<\/em>.<\/p>\n<p>\ud83d\udca1 <strong>Connecting the Dots:<\/strong> o ativo de autoridade aqui n\u00e3o \u00e9 a Average Precision de 0,60 \u2014 \u00e9 o fato de que <strong>a AUROC \u00e9 matematicamente insens\u00edvel \u00e0 preval\u00eancia<\/strong>, e portanto o mesmo 0,95 descreve um modelo com VPP de 90% e outro com VPP de 10%. Todo mundo cita a AUROC; quase ningu\u00e9m no meio m\u00e9dico brasileiro traduz que <em>ela n\u00e3o sabe se o desfecho \u00e9 raro, e por isso n\u00e3o sabe nada sobre o que vai acontecer no plant\u00e3o<\/em>. Mas o segundo salto \u00e9 o que quase ningu\u00e9m d\u00e1: se sensibilidade alta com VPP baixo \u00e9 a assinatura inevit\u00e1vel do desfecho raro, ent\u00e3o a vari\u00e1vel de projeto n\u00e3o \u00e9 o algoritmo \u2014 \u00e9 <strong>o ato cl\u00ednico que a sa\u00edda solicita<\/strong>. &#8220;Risco alto&#8221; exige VPP que o desfecho raro n\u00e3o permite entregar; &#8220;reavalia\u00e7\u00e3o em 15 minutos&#8221; n\u00e3o exige. E como a CFM 2.454 classifica por <em>grau de autonomia<\/em> e n\u00e3o por acur\u00e1cia, reescrever aquela frase na tela reduz simultaneamente a exig\u00eancia estat\u00edstica e a classe de risco regulat\u00f3ria. \u00c9 engenharia cl\u00ednica e engenharia regulat\u00f3ria sendo a mesma decis\u00e3o \u2014 e essa decis\u00e3o n\u00e3o se aprende em curso de machine learning, porque exige saber o que acontece com a enfermagem no terceiro m\u00eas, nem em resid\u00eancia, porque exige saber por que a curva precis\u00e3o-recall existe. \u00c9 exatamente essa interse\u00e7\u00e3o, e n\u00e3o o acesso ao modelo, que constitui o fosso t\u00e9cnico de quem constr\u00f3i IA cl\u00ednica no Brasil.<\/p>\n<\/div>\n<\/section>\n<section class=\"section\" id=\"ref-pt\">\n<div class=\"reference-box\">\n<h4>Refer\u00eancias<\/h4>\n<ol>\n<li>Ha T, Kappy B, Chamberlain JM, McKinley KW. <em>Early Prediction of Critical Care Interventions From Pediatric Emergency Department Triage<\/em>. Hosp Pediatr. 2026 Ago;16(8):e618\u2013e625. (886.183 visitas; 26.721 desfechos, 3,0%; rede neural AP 0,60 IC95% 0,59\u20130,61; sensibilidade 88%, VPP 32%; contrafactual na Tabela 4) Dispon\u00edvel em: https:\/\/doi.org\/10.1542\/hpeds.2025-009127<\/li>\n<li>Davis J, Goadrich M. <em>The Relationship Between Precision-Recall and ROC Curves<\/em>. Proceedings of the 23rd International Conference on Machine Learning (ICML). 2006:233\u2013240. Dispon\u00edvel em: https:\/\/doi.org\/10.1145\/1143844.1143874<\/li>\n<li>Saito T, Rehmsmeier M. <em>The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets<\/em>. PLoS One. 2015;10(3):e0118432. (acesso aberto; ver p. 5 para a linha de base da PRC, y = P\/(P+N), e Fig 2B para a demonstra\u00e7\u00e3o visual) Dispon\u00edvel em: https:\/\/doi.org\/10.1371\/journal.pone.0118432<\/li>\n<li>Conselho Federal de Medicina. <em>Resolu\u00e7\u00e3o CFM n\u00ba 2.454, de 11 de fevereiro de 2026 \u2014 Normatiza o uso da intelig\u00eancia artificial na medicina<\/em>. DOU 2026 fev 27; Ed. 39, Se\u00e7\u00e3o 1. Dispon\u00edvel em: https:\/\/sistemas.cfm.org.br\/normas\/arquivos\/resolucoes\/BR\/2026\/2454_2026.pdf<\/li>\n<li>Jiang X, Yang N, Shi T, et al. <em>Deciphering the &#8220;non-verbal&#8221; code: A preliminary exploration of multimodal large language models for neonatal pain recognition<\/em>. Digit Health. 2026;12:20552076261473729. (mesma assimetria sensibilidade\/especificidade em dom\u00ednio distinto) Dispon\u00edvel em: https:\/\/doi.org\/10.1177\/20552076261473729<\/li>\n<li>Lonsdale H, Patel K, Domenico H, et al. <em>Development and external validation of the NEO-READY model to predict date of discharge among premature neonatal intensive care patients<\/em>. J Perinatol. 2026. Dispon\u00edvel em: https:\/\/doi.org\/10.1038\/s41372-026-02827-2<\/li>\n<li><em>Artificial Intelligence in Pediatric Cardiac Intensive Care: Clinical Applications, Implementation Challenges, and Future Directions<\/em>. Curr Treat Options Pediatr. 2026. Dispon\u00edvel em: https:\/\/link.springer.com\/article\/10.1007\/s40746-026-00374-8<\/li>\n<\/ol>\n<\/div>\n<\/section>\n<\/div>\n<\/div>\n<p><!-- ===================== CONTE\u00daDO EN ===================== --><\/p>\n<div id=\"content-en\" class=\"hidden\">\n<section class=\"hero\" id=\"top-en\">\n<span class=\"article-tag\">Pediatric Emergency \u00b7 Methodology \u00b7 Machine Learning<\/span><\/p>\n<h1>The AI That Finds the Critically Ill Child in the ED Queue \u2014 and the Metric That Makes Every Model Look Better Than It Is<\/h1>\n<p class=\"subtitle\">A model trained on 886,000 pediatric encounters identifies, using triage data alone, 88% of the children who will need a critical care intervention \u2014 and would triple the share of intermediate-risk patients evaluated in time. That is headline material. But there is a fact almost no clinical article translates: <strong>the metric that made these models famous is blind to what matters most in triage \u2014 how rare the event is<\/strong>. Understanding why AUROC lies for rare outcomes, and what to put in its place, is what separates someone who consumes a predictive score from someone who can judge one.<\/p>\n<p class=\"meta-info\">Dr. Mbula Luzingu Barros | Pediatric Intensivist \u00b7 25 years in Pediatric ICU | AI Healthcare Consultant | Founder of INOVAMED | Creator of the AIMED Methodology<\/p>\n<p class=\"date\">\ud83d\udcc5 Published July 27, 2026<\/p>\n<div class=\"deadline-badge\">\u23f1 CFM 2.454 full enforcement in: <span class=\"countdown-days\">&#8212;<\/span> days \u2014 August 26, 2026<\/div>\n<\/section>\n<div class=\"container\">\n<div class=\"toc\" id=\"indice-en\">\n<h2>Table of Contents<\/h2>\n<p class=\"toc-subtitle\">From the emergency queue to the regulatory risk class \u2014 why choosing the metric is choosing the clinical outcome<\/p>\n<ul>\n<li><a href=\"#s1-en\">1. Why This Matters for Those on the Front Line<\/a><\/li>\n<li><a href=\"#s2-en\">2. What the Study Delivered<\/a><\/li>\n<li><a href=\"#s3-en\">3. The Numbers, Unfiltered<\/a><\/li>\n<li><a href=\"#s4-en\">4. The Trap: The Metric That Cannot Feel Rarity<\/a><\/li>\n<li><a href=\"#s5-en\">5. Choosing the Metric Is a Clinical Act<\/a><\/li>\n<li><a href=\"#s6-en\">6. The Regulatory Layer \u2014 CFM 2.454<\/a><\/li>\n<li><a href=\"#s7-en\">7. The Timeline<\/a><\/li>\n<li><a href=\"#s8-en\">8. What to Do on Monday<\/a><\/li>\n<li><a href=\"#final-en\">9. Final Considerations<\/a><\/li>\n<li><a href=\"#ref-en\">10. References<\/a><\/li>\n<\/ul>\n<\/div>\n<div class=\"science-box\">\n<h4>For readers outside healthcare \u2014 the essentials in 60 seconds<\/h4>\n<p>This article works on two layers. The technical one, for people who work in emergency departments or build models; and the general one, which requires only curiosity. Every core idea appears twice: once with a hospital example, once with an everyday one. If the first does not land, skip to the second \u2014 you will not miss anything.<\/p>\n<p><strong>Triage<\/strong> is what happens when you arrive at an emergency department: someone decides who is seen first. Not order of arrival, but order of severity. Brazil uses the <strong>Manchester Protocol<\/strong>, which sorts patients into five colors \u2014 red is now, yellow is soon, green can wait. The United States uses the <strong>ESI<\/strong>, which does the same with numbers 1 to 5.<\/p>\n<p><strong>The respiratory support acronyms<\/strong> ahead are simply degrees of help with breathing, from lightest to heaviest: <em>HFNC<\/em> (high-flow oxygen through the nose) and continuous nebulization are the light ones; <em>CPAP<\/em> and <em>BiPAP<\/em> are masks that push air under pressure; <em>intubation<\/em> is the tube in the windpipe, with a machine breathing for the patient. A <em>vasopressor<\/em> is the medication that holds up blood pressure when the body can no longer do it alone.<\/p>\n<p><strong>Prevalence<\/strong> is the fraction of people in whom something happens. &#8220;A prevalence of 3%&#8221; means that, of every 100 encounters, 3 end in that outcome. It is the most important word in this article \u2014 keep hold of it.<\/p>\n<\/div>\n<section class=\"section\" id=\"s1-en\">\n<h2 class=\"section-title\">Why This Matters for Those on the Front Line<\/h2>\n<div class=\"intro-box\">\n<p>Every pediatric triage system has a purgatory. In the American system it is <strong>ESI level 3<\/strong>; in the Manchester system, which is ours in Brazil, it is <strong>yellow<\/strong>. It is the band where the child is not obvious enough to be red and not well enough to be green. It is where the bronchiolitis that will tire out in four hours lives, alongside the dehydration that has not yet closed the capillary refill, and the early sepsis that still has only a fever and nothing else.<\/p>\n<p>The problem with tiered risk classification is not that it errs \u2014 it is that it <strong>saturates<\/strong>. A five-level system has, by construction, five degrees of resolution; and most of the movement happens inside a single level. The triage nurse knows this. The attending knows this. And that is exactly the gap a predictive model proposes to fill: not to replace Manchester, but to <strong>order what is stacked inside the yellow<\/strong>.<\/p>\n<p>A study published on July 20, 2026 did precisely that, at scale. And the technical answer it gave is good. But what makes this paper worth a critical read is not the model \u2014 it is the <em>metric<\/em> the authors chose to report it with, and why that choice is itself a clinical decision.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s2-en\">\n<h2 class=\"section-title\">What the Study Delivered<\/h2>\n<div class=\"detailed-section\">\n<p>A retrospective cohort from 2016 to 2024, at a large urban academic pediatric emergency department in the United States with roughly 85,000 annual encounters. <strong>886,183 encounters.<\/strong> The outcome occurred in <strong>26,721 visits, or 3.0% of the total<\/strong>.<\/p>\n<h4>What exactly counts as the outcome \u2014 read this before comparing to your unit<\/h4>\n<p>The authors defined &#8220;need for critical care interventions&#8221; as the <strong>use of life-saving medications or respiratory support within 48 hours of ED presentation<\/strong>. Medications: continuous infusion of vasopressors, insulin, or terbutaline. Respiratory support: intubation, BiPAP, CPAP, heliox, <strong>high-flow nasal cannula (HFNC)<\/strong>, and <strong>at least 2 hours of continuous nebulized albuterol<\/strong>. It also counts what was done on an inpatient unit and what was done at a return ED visit within 48 hours.<\/p>\n<p>That is broader than it sounds. HFNC and two hours of continuous albuterol happen routinely in the observation area and on the ward \u2014 they are not, to a pediatric intensivist, &#8220;critical care.&#8221; <strong>A good share of that 3% is asthma and bronchiolitis receiving standard treatment, not a child on the verge of collapse.<\/strong> When you read the number, adjust the expectation: the model predicts the need to escalate support, not imminent arrest.<\/p>\n<p>Here the authors did the right thing and deserve credit: they ran a <strong>more stringent secondary outcome<\/strong>, restricted to intubation, BiPAP, CPAP, and heliox \u2014 excluding HFNC and continuous albuterol \u2014 and performance held. They also tested a shorter 8-hour window, with similar results. Those are two sensitivity analyses most papers skip.<\/p>\n<p>The authors trained six algorithms using <strong>exclusively information available at the moment of triage<\/strong>. No laboratory results, no evolution, no reassessment \u2014 only what the triage nurse holds in the first minutes. The neural network performed best.<\/p>\n<h4>The number that matters is not the model&#8217;s<\/h4>\n<p>The authors ran a counterfactual simulation: what would happen to time-to-physician if the score were used <em>alongside<\/em> the ESI rather than in place of it? In the ESI 3 band \u2014 the purgatory \u2014 the proportion of critical care patients evaluated in time would jump from <strong>23.3% to 75.0%<\/strong>, and median time-to-pediatrician would fall from <strong>34 to 10 minutes<\/strong>. For ESI 2, from 48.7% to 87.1% (19 \u2192 12 min). For ESI 4, from 10.3% to 60.7% (62 \u2192 7 min).<\/p>\n<p>Note the design: the model reclassifies no one. It <em>reorders within the class<\/em>. That is an architectural choice that preserves the validated instrument the team already uses and adds resolution exactly where that instrument is blind. It is worth more as a lesson in clinical engineering than the AUC of any of the six algorithms.<\/p>\n<\/div>\n<div class=\"highlight-box\">\n<h4>A concrete example: what &#8220;predicting&#8221; means here<\/h4>\n<p>It is 9 p.m. on a winter Saturday. Reception holds 40 children waiting, 26 of them classified yellow. Manchester has already done its job: it separated the 3 reds and the 11 greens. What remains are 26 yellows, mutually indistinguishable on the screen, ordered by arrival time.<\/p>\n<p>The model changes no one&#8217;s color. It reorders the 26 \u2014 and puts at the top the 5-month-old infant with a respiratory rate in the high percentile for age and a saturation of 93%, who arrived 40 minutes after a teenager with abdominal pain. <strong>None of that information is new.<\/strong> All of it was on the triage form. The model merely did the arithmetic nobody has time to do at 9 p.m. on a Saturday.<\/p>\n<\/div>\n<div class=\"warning-box\">\n<h4>\u26a0 The counterfactual is the weakest link \u2014 and it is worth understanding why<\/h4>\n<p>The simulation rule is this: each patient who <em>actually<\/em> received a critical care intervention is assigned the earliest evaluation timestamp among all <em>non-critical<\/em> patients of the same ESI waiting concurrently. &#8220;Timely evaluation&#8221; means having been seen ahead of every non-critical patient at that same level.<\/p>\n<p>Note what the rule does <strong>not<\/strong> model: the <em>false positives<\/em>. With a PPV of 32%, for every 1,000 encounters the model raises roughly 82 flags and only 26 are real \u2014 the other 56 are children who would also be pushed to the front of the queue, competing for exactly the same priority slots. The queue has finite capacity: there is only one &#8220;next available pediatrician.&#8221; The simulation grants the benefit to the true positives without charging them the competition from the false ones.<\/p>\n<p>There is a productive irony here, and it is not in the paper&#8217;s limitations section: <strong>the same work that honestly reports a PPV of 32% runs a counterfactual that behaves, in practice, as if the PPV were far higher.<\/strong> This does not invalidate the study \u2014 but it turns the 23.3% \u2192 75.0% into the <em>ceiling<\/em> of what the information could buy, not what implementation would deliver. Add to that what the authors themselves concede: single center, no prospective validation, with a declared need for recalibration at other sites.<\/p>\n<\/div>\n<div class=\"alert-box\">\n<h4>\u26a0 The transfer blocker to Brazil that almost no one will notice<\/h4>\n<p>It sits in a single sentence of the limitations section: <em>&#8220;our models depend on natural language processing of nursing narratives.&#8221;<\/em> The model does not run on structured vital signs \u2014 it reads the <strong>free text the nurse writes at triage<\/strong>.<\/p>\n<p>That changes everything for us. First, it is NLP trained on English, over American nursing documentation conventions. Second, and more serious: Brazilian Manchester triage is heavily structured around discriminators, and the quality and volume of free text vary enormously between services \u2014 in many, it is one line. <strong>The variable carrying the most signal in the original model is precisely the one that barely exists in the Brazilian workflow.<\/strong> This is not a recalibration case: it is a case for retraining on a different feature base.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s3-en\">\n<h2 class=\"section-title\">The Numbers, Unfiltered<\/h2>\n<p>Before the table, five concepts. They sound dry, but each answers a question you already ask in your head on call \u2014 just without the technical name. Worth reading even if statistics is not your terrain: this article&#8217;s entire argument fits here.<\/p>\n<p><em>A note on the &#8220;95% CI&#8221; appearing in the table: it is the 95% confidence interval \u2014 the range within which the true value probably lies. A narrow range, like the study&#8217;s 0.59\u20130.61, means a precise estimate, a consequence of the sample being enormous.<\/em><\/p>\n<div class=\"science-box\">\n<h4>The metric decoder (start with the threshold \u2014 everything depends on it)<\/h4>\n<p><strong>Threshold \u2014 &#8220;from what point do I call?&#8221;<\/strong> A model does not return &#8220;yes&#8221; or &#8220;no.&#8221; It returns a continuous number, something like 0.17 or 0.64. Someone has to decide from which value that becomes an alert on the screen. That cut-off point is the <strong>threshold<\/strong>.<\/p>\n<p><em>In the hospital:<\/em> the saturation at which you decide to call the attending is a threshold. Raise it from 92% to 94% and you call earlier and more often: you catch nearly every child who was going to worsen, and you call a lot of people for nothing. Drop it to 88% and you call less often and are almost always right \u2014 but you miss the ones who were just starting to sink.<\/p>\n<p><em>Outside it:<\/em> it is the strictness setting on your spam filter. Too tight, and real email lands in the trash. Too loose, and junk floods your inbox. You cannot have both \u2014 and no adjustment of the filter fixes it, because the problem is not the filter, it is having to pick one point on a single scale.<\/p>\n<p><strong>Sensitivity and positive predictive value are exactly that balance, and moving the threshold trades one for the other.<\/strong> There is no setting that improves both at once: it is the same rope, pulled from opposite ends. Hold on to this \u2014 it is the piece that carries this article&#8217;s conclusion.<\/p>\n<p><strong>Sensitivity \u2014 &#8220;of those who would need the ICU, how many does the model catch?&#8221;<\/strong> Sensitivity of 88% means that, of 100 children who will receive a critical care intervention, the model raises the alert on 88. The other 12 slip through. It is the metric that hurts to miss.<\/p>\n<p><strong>Positive predictive value (PPV) \u2014 &#8220;when it calls me, how often is it real?&#8221;<\/strong> A PPV of 32% means that, of every 100 alerts fired, 32 are children who really will need care. The other 68 are not. That is the number that decides whether the team is still looking at the alert in week three.<\/p>\n<p><strong>AUROC \u2014 &#8220;can it tell which of two children is sicker?&#8221;<\/strong> Draw two children at random, one who will deteriorate and one who will not. The AUROC is the probability the model gives the higher score to the one who deteriorates. It is a test of <em>pairwise ranking<\/em>. And here is the key this entire article pursues: <strong>that test does not change whether the event is common or vanishingly rare<\/strong>. AUROC is mathematically insensitive to prevalence.<\/p>\n<p><strong>Average Precision (AP) \u2014 &#8220;of the times it called me, how many were worth it?&#8221;<\/strong> It is the mean PPV, computed by sweeping every possible threshold at once. Instead of giving you the PPV at one specific cut-off, it gives you the behavior of the whole model. And, unlike AUROC, AP <em>feels<\/em> how rare the event is.<\/p>\n<p><strong>The &#8220;baseline&#8221; of a metric<\/strong> is the score a <em>useless<\/em> model earns \u2014 the one flipping a coin. It works as a ruler: without knowing the useless model&#8217;s score, you cannot tell whether the good model&#8217;s score is good. And here is the difference that decides everything, demonstrated by Saito and Rehmsmeier:<\/p>\n<p>The <strong>AUROC baseline is always 0.50<\/strong>, whether the outcome happens in half of patients or in one per thousand. The <strong>AP baseline is the prevalence itself<\/strong> \u2014 the authors show it equals exactly P\/(P+N), the fraction of positive cases in the total. In plain shift-floor terms: for an outcome occurring in 3% of children, flipping a coin earns an AP of <strong>0.03<\/strong>. Not 0.50 \u2014 0.03.<\/p>\n<p>That is why the study&#8217;s AP of 0.60 means <strong>twenty times better than chance<\/strong>. Had it been an AUROC of 0.60, it would be barely above a coin flip. Same number, opposite readings \u2014 because the rulers are different.<\/p>\n<\/div>\n<div class=\"table-wrapper\">\n<table class=\"data-table\">\n<thead>\n<tr>\n<th>Metric<\/th>\n<th>Reported value<\/th>\n<th>Baseline<\/th>\n<th>Source<\/th>\n<th>Reading<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Cohort size<\/td>\n<td>886,183 encounters (2016\u20132024)<\/td>\n<td>\u2014<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-baixo\">\u2191 robust<\/span><\/td>\n<\/tr>\n<tr>\n<td>Outcome prevalence<\/td>\n<td>26,721 visits \u00b7 <strong>3.0%<\/strong><\/td>\n<td>\u2014<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-alto\">\u26a0 rare event<\/span><\/td>\n<\/tr>\n<tr>\n<td><strong>Average Precision<\/strong> (neural network)<\/td>\n<td><strong>0.60<\/strong> (95% CI 0.59\u20130.61)<\/td>\n<td>0.03<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-baixo\">\u2191 20\u00d7 chance<\/span><\/td>\n<\/tr>\n<tr>\n<td>Sensitivity<\/td>\n<td>88% (87\u201389%)<\/td>\n<td>\u2014<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-medio\">\u2191 good<\/span><\/td>\n<\/tr>\n<tr>\n<td>Positive predictive value<\/td>\n<td>32% (31\u201332%)<\/td>\n<td>3%<\/td>\n<td>Ha et al. \u00b7 Hosp Pediatr \u00b7 2026<\/td>\n<td><span class=\"status-badge s-alto\">\u2193 2 in 3 alerts are false<\/span><\/td>\n<\/tr>\n<tr>\n<td>Specificity <em>(derived, not reported)<\/em><\/td>\n<td>\u2248 94%<\/td>\n<td>\u2014<\/td>\n<td>Own calculation from sens. and PPV<\/td>\n<td><span class=\"status-badge s-inaceitavel\">\u2192 see box below<\/span><\/td>\n<\/tr>\n<tr>\n<td>ESI 3 evaluated in time<\/td>\n<td>23.3% \u2192 <strong>75.0%<\/strong><\/td>\n<td>\u2014<\/td>\n<td>Ha et al. \u00b7 counterfactual analysis<\/td>\n<td><span class=\"status-badge s-medio\">\u2191 ceiling, not floor<\/span><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<div class=\"science-box\">\n<h4>The arithmetic that reorganizes the whole table<\/h4>\n<p>Specificity was not reported in the study abstract, but it can be derived. With a prevalence of 3%, sensitivity of 88%, and PPV of 32%, for every <strong>1,000 encounters<\/strong>: 30 children will need an intervention; the model catches 26 of them; to reach a PPV of 32% it must fire roughly <strong>82 alerts in total<\/strong>; therefore 56 are false positives. That leaves 914 non-events correctly kept silent out of 970. <strong>Specificity \u2248 94%.<\/strong><\/p>\n<p>Now read the following two sentences, which describe exactly the same classifier: <em>&#8220;specificity of 94%&#8221;<\/em> and <em>&#8220;two out of every three alerts are false.&#8221;<\/em> The first passes any committee. The second describes the nursing staff&#8217;s actual experience in month three. <strong>There is no contradiction \u2014 there are two different questions.<\/strong> Specificity asks what happens to those who are fine; PPV asks what happens to those who were called. For rare events the denominators are brutally different in size, which is why the two answers diverge so sharply.<\/p>\n<p><strong>If you have ever been through an airport, you have lived this.<\/strong> The metal detector catches essentially every weapon that passes through it \u2014 very high sensitivity, exactly as it should be. It also stays silent for the overwhelming majority of passengers \u2014 very high specificity. And still, when it beeps, it is almost never a weapon: it is a belt buckle, a key, a coin forgotten in a pocket. <strong>All three things are true at once<\/strong>, and they only seem contradictory until you notice that weapons are vanishingly rare among passengers. That is exactly the structure of this article&#8217;s problem \u2014 and the reason nobody switches off the airport detector while everybody switches off the hospital alert: at the airport, the cost of a false alarm is ten seconds with a wand. In the hospital, when the cost of a false alarm is high, the alert dies.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s4-en\">\n<h2 class=\"section-title\">The Trap: The Metric That Cannot Feel Rarity<\/h2>\n<div class=\"detailed-section\">\n<p>Here is the fact almost no clinical popularizer translates: <strong>AUROC does not change when prevalence changes.<\/strong> This is neither a practical limitation nor a sampling bias \u2014 it is a mathematical property. AUROC is built on sensitivity and specificity, and both are computed <em>within<\/em> each group: sensitivity looks only at the sick, specificity only at the well. Neither knows the ratio between the groups.<\/p>\n<p>The consequence is uncomfortable. <strong>The same AUROC value describes a model with 90% PPV and a model with 10% PPV<\/strong>, depending solely on how rare the outcome is. If you read &#8220;AUROC 0.95&#8221; without knowing the prevalence, you know absolutely nothing about how the alert will behave on shift.<\/p>\n<h4>Why this is epidemic in pediatric literature<\/h4>\n<p>Because nearly every serious pediatric outcome is rare. Cardiac arrest, sepsis, need for ventilation, death \u2014 all with single-digit prevalence. That is precisely the regime in which AUROC is most generous and least informative. A mediocre model on a 2%-prevalence outcome easily produces an AUROC of 0.90, because it is easy to separate the overwhelming majority of well children from the few sick ones; the hard part is getting <em>which<\/em> of the few right.<\/p>\n<p>This critique is neither mine nor new. Davis and Goadrich formally demonstrated, in 2006, that the precision-recall curve is more informative on <em>imbalanced<\/em> datasets \u2014 jargon for &#8220;almost everyone is negative and very few are positive,&#8221; which is the exact description of any serious pediatric outcome \u2014 and that optimizing the area under the ROC does not guarantee optimizing the area under the precision-recall curve. Saito and Rehmsmeier, in 2015, put it plainly in their own abstract: the visual interpretability of ROC plots on imbalanced datasets <em>&#8220;can be deceptive with respect to conclusions about the reliability of classification performance, owing to an intuitive but wrong interpretation of specificity.&#8221;<\/em> <strong>Twenty years of established methodology, nearly 4,700 citations, and pediatric clinical literature still reports AUROC for 3%-prevalence outcomes.<\/strong><\/p>\n<\/div>\n<div class=\"warning-box\">\n<h4>\u26a0 Where my own argument has limits: prevalence is not the same as spectrum<\/h4>\n<p>I need to be precise here, because an attentive reader will find this on their own. Saying &#8220;AUROC is insensitive to prevalence&#8221; is a statement about the <strong>construction of the metric<\/strong>: sensitivity and specificity are ratios computed within each group, which is why the ROC baseline does not move when the proportion between groups changes.<\/p>\n<p>That does <em>not<\/em> mean the measured AUROC will be the same in any service. If the new population has sicker cases, milder cases, or a different mix, the scores the model hands out spread differently \u2014 and the AUROC you measure in practice shifts with them. That is the <strong>spectrum effect<\/strong>: the proportion of sick patients did not change, the <em>kind<\/em> of sick patient did. These are two distinct problems that often appear together and get conflated: <em>prevalence invariance<\/em> is a mathematical property of the metric; the <em>spectrum effect<\/em> is a change in the population.<\/p>\n<p>The study under discussion illustrates the second problem well: the cohort is 61.8% non-Hispanic Black children and 23.2% Hispanic or Latino, with 74% covered by public insurance, at a single hospital in Washington, DC. It is a real and well-described population \u2014 and it is <em>another<\/em> population.<\/p>\n<p>The practical consequence, however, is the same in both cases, and it reinforces rather than weakens the argument: <strong>a published AUROC does not let you anticipate how many false alerts your team will receive.<\/strong> In the first case because the metric does not carry that information; in the second because the population is not the same. Either way, the number you need is the PPV computed at your prevalence.<\/p>\n<\/div>\n<div class=\"alert-box\">\n<h4>\u26a0 The risk is not the isolated false positive \u2014 it is the death of the alert<\/h4>\n<p>Two false alerts out of three are not a statistical inconvenience. They are the mechanism of <em>alarm fatigue<\/em>, and alarm fatigue has documented clinical consequences: an alert nobody looks at protects less than no alert at all, because it consumes attention and produces a false sense of coverage. A predictive triage system with low PPV and the wrong presentation does not degrade gracefully \u2014 it is mentally switched off by the team within weeks, and stays switched on on the screen for months.<\/p>\n<\/div>\n<div class=\"highlight-box\">\n<h4>A concrete example: the saturation alarm you already silenced today<\/h4>\n<p>You know this metric without knowing its name. The pulse oximeter on bed 4 beeps every twenty minutes. The overwhelming majority of the time it is a displaced sensor, a child who moved a hand, an artifact. That alarm&#8217;s <em>specificity<\/em> is very high \u2014 it stays quiet most of the time, relative to all the minutes the child is fine. Its <em>positive predictive value<\/em> is very low \u2014 when it screams, it is almost never a real desaturation.<\/p>\n<p><em>Outside the hospital, the same:<\/em> think of the car alarm that goes off in the street at three in the morning. Nobody gets up. Nobody looks out the window. Not because the alarm is never right \u2014 but because it is almost always wrong, and the human brain learns that within weeks. A car alarm has excellent sensitivity for break-ins and negligible positive predictive value, which is why it became urban noise instead of a security system.<\/p>\n<p>You do not silence the oximeter because specificity is poor. You silence it because PPV is poor. <strong>Anyone already knows intuitively which of the two metrics governs human behavior \u2014 they just do not know it has a name.<\/strong> What is missing is demanding that papers report the very one everyone already uses in practice.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s5-en\">\n<h2 class=\"section-title\">Choosing the Metric Is a Clinical Act<\/h2>\n<div class=\"detailed-section\">\n<p>The real merit of Ha and colleagues&#8217; study is not the neural network. It is that they reported <strong>Average Precision<\/strong> as the primary metric for a 3%-prevalence outcome, and reported the sensitivity\/PPV pair instead of sensitivity\/specificity. That is rare, and it is why their reading can be trusted.<\/p>\n<p>But the choice of metric is not a reporting detail. It encodes a decision about <strong>what the model is<\/strong>. Optimizing sensitivity produces a screening instrument. Optimizing PPV produces a decision instrument. These are different clinical products, with different workflows and different responsibilities \u2014 and the literature treats the choice as if it were statistical preference.<\/p>\n<h4>The architectural consequence almost nobody draws<\/h4>\n<p>Sensitivity of 88% with a PPV of 32% is not a badly tuned classifier. It is the <strong>signature<\/strong> of any model optimized on an outcome that is rare and costly to miss. You do not fix that by moving the threshold \u2014 moving the threshold merely trades one metric for the other along the same curve.<\/p>\n<p>The way out is structural: a <strong>two-stage cascade<\/strong>. Stage 1 cheap and sensitive, sweeping everyone and accepting false positives. Stage 2 expensive and specific, triggered <em>only<\/em> on stage 1&#8217;s positives \u2014 and stage 2 can perfectly well be human. A five-minute structured nursing reassessment is a high-specificity classifier that already exists, is already validated, and is already in the service.<\/p>\n<\/div>\n<div class=\"highlight-box\">\n<h4>A concrete example: verdict versus summons<\/h4>\n<p>There are two ways the same model, with exactly the same statistical performance, can appear on the triage screen.<\/p>\n<p><strong>As a verdict:<\/strong> <em>&#8220;High risk of needing critical care.&#8221;<\/em> For that to be accepted, PPV must be high \u2014 otherwise the team discovers within three weeks that the machine is wrong two times out of three and stops looking. With a PPV of 32%, that product dies.<\/p>\n<p><strong>As a summons:<\/strong> <em>&#8220;Structured reassessment suggested within 15 minutes.&#8221;<\/em> For that to be accepted, PPV <em>need not<\/em> be high \u2014 it only requires that the cost of reassessment be low. Five minutes of nursing time, 82 times per thousand encounters. With a PPV of 32%, that product works.<\/p>\n<p><strong>Same model. Same numbers. One dies, the other lives.<\/strong> The difference is not in the algorithm \u2014 it is in which clinical act the output requests.<\/p>\n<\/div>\n<div class=\"warning-box\">\n<h4>\u26a0 Where the cascade can still fail<\/h4>\n<p>The cascade only works if stage 2 is genuinely more specific than stage 1 <em>and<\/em> carries low marginal cost. If the structured reassessment is not protocolized, what the summons produces is 82 interruptions per thousand encounters with no information gain \u2014 the same alarm fatigue, under a different label. Stage 2 is the product; the model is only the trigger.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s6-en\">\n<h2 class=\"section-title\">The Regulatory Layer \u2014 CFM 2.454<\/h2>\n<div class=\"detailed-section\">\n<p>CFM Resolution No. 2,454\/2026 classifies AI systems by risk level \u2014 low, medium, high, or unacceptable \u2014 considering impact on fundamental rights, model complexity, <strong>degree of autonomy<\/strong>, and data sensitivity. The 180-day adaptation period ends on <strong>August 26, 2026<\/strong>, and the duty to classify falls on the <em>implementing institution<\/em>, not the vendor.<\/p>\n<h4>The detail that closes this article&#8217;s argument<\/h4>\n<p>Reread the criteria and note the second one: <strong>degree of autonomy<\/strong>. A system issuing a priority verdict about a child in an emergency queue has high decisional autonomy and direct impact on access to care \u2014 it hardly escapes high risk, with all the burden that carries: documented validation, physician oversight of the output, and a chart entry recording AI support for the decision.<\/p>\n<p>A system that merely <em>summons a human reassessment<\/em> has substantially lower decisional autonomy: it decides nothing, it schedules a look. The plausible classification drops to medium, possibly low. <strong>The same model, with the same performance, changes regulatory class according to the clinical act it requests.<\/strong> This is not a loophole \u2014 it is the norm working as intended: it regulates autonomy, not accuracy.<\/p>\n<\/div>\n<div class=\"alert-box\">\n<h4>\u26a0 What is in use today and nobody has classified<\/h4>\n<p>The resolution reaches <strong>tools already in operation<\/strong>. That includes the category no service is mapping: the general-purpose LLM accessed informally by team members within the care workflow. On August 26, that is an unclassified AI system inside a medical institution. The defensible response is not to ban it \u2014 it is to inventory and classify it before the deadline.<\/p>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s7-en\">\n<h2 class=\"section-title\">The Timeline<\/h2>\n<div class=\"flow-box\">\n<h4>From the published metric to the risk class in your institution<\/h4>\n<p class=\"flow-intro\">Why the decision that determines a pediatric predictive model&#8217;s fate is made before the first line of code \u2014 and after the last.<\/p>\n<div class=\"flow-steps\">\n<div class=\"flow-step\">\n<div class=\"step-label\">Publication<\/div>\n<div class=\"step-title\">High AUROC on a rare event<\/div>\n<div class=\"step-desc\">The prevalence-insensitive metric produces the headline. PPV does not appear in the abstract.<\/div>\n<\/div>\n<div class=\"flow-arrow\">\u279c<\/div>\n<div class=\"flow-step\">\n<div class=\"step-label\">Reality<\/div>\n<div class=\"step-title\">Two in three alerts are false<\/div>\n<div class=\"step-desc\">94% specificity and 32% PPV coexist. The team silences the alert within weeks.<\/div>\n<\/div>\n<div class=\"flow-arrow\">\u279c<\/div>\n<div class=\"flow-step current\">\n<div class=\"step-label\">Today \u00b7 You are here<\/div>\n<div class=\"step-title\">Reframe the clinical act<\/div>\n<div class=\"step-desc\">A reassessment summons instead of a priority verdict. Two-stage cascade.<\/div>\n<\/div>\n<div class=\"flow-arrow\">\u279c<\/div>\n<div class=\"flow-step\">\n<div class=\"step-label\">Aug 26, 2026<\/div>\n<div class=\"step-title\">CFM 2.454 classification<\/div>\n<div class=\"step-desc\">Lower autonomy, lower risk class, lower compliance burden. The institution&#8217;s duty.<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/section>\n<div class=\"divider\"><\/div>\n<section class=\"section\" id=\"s8-en\">\n<h2 class=\"section-title\">What to Do on Monday<\/h2>\n<div class=\"highlight-box\">\n<h4>From critically reading a paper to classifying your own system<\/h4>\n<div class=\"checklist-section\">\n<div class=\"check-subhead\">When reading any pediatric predictive model paper<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Find the outcome prevalence before looking at the AUROC.<\/strong> Below 10%, treat AUROC alone as uninformative and demand AUPRC or Average Precision.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Check that AP comes with its baseline.<\/strong> AP without the prevalence beside it is a number without a scale \u2014 the AP baseline <em>is<\/em> the prevalence.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>If only sensitivity and specificity are given, derive the PPV yourself.<\/strong> All you need is the prevalence. It is the calculation that reveals how many false alerts the team will receive.<\/div>\n<\/div>\n<div class=\"check-subhead\">Before putting any model into production<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Decide whether the product is screening or decision<\/strong> \u2014 and report the matching metric. Screening answers to sensitivity; decision answers to PPV.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Design stage 2 before stage 1.<\/strong> If the structured reassessment is not protocolized and timed, the model only produces interruption.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Write the exact text that appears on screen.<\/strong> &#8220;High risk&#8221; and &#8220;reassessment suggested within 15 minutes&#8221; carry different PPV requirements and different regulatory risk classes.<\/div>\n<\/div>\n<div class=\"check-subhead\">Before August 26<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Inventory the AI systems in use in your unit<\/strong> across four categories: embedded in equipment, contracted diagnostic software, general-purpose LLMs accessed by the team, and internally developed models.<\/div>\n<\/div>\n<div class=\"check-item\"><span class=\"check-icon\">\u2610<\/span><\/p>\n<div class=\"check-text\"><strong>Classify each by degree of autonomy<\/strong>, not by accuracy. That is the norm&#8217;s criterion.<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/section>\n<div class=\"cta-section\">\n<h2>Learn to Judge a Model Before Trusting It<\/h2>\n<p>The AIMED methodology trains physicians who build \u2014 not merely consume \u2014 clinical AI tools. Metric selection under class imbalance, cascade architecture, clinical act design, and CFM 2.454 risk classification are part of the curriculum, because they are the same decision seen from different angles.<\/p>\n<p><a href=\"https:\/\/inovamed.pro\/?page_id=96\" class=\"cta-button\">Discover AIMED \u2192<\/a>\n<\/div>\n<section class=\"section\" id=\"final-en\">\n<h2 class=\"section-title\">Final Considerations<\/h2>\n<div class=\"intro-box\">\n<p>Ha and colleagues&#8217; model is good. An Average Precision of 0.60 against a baseline of 0.03 is twenty times chance, across 886,000 encounters, using triage data alone. There is nothing to diminish in the work \u2014 on the contrary, it is exemplary precisely because it reported the hard metric when it could have reported the easy one.<\/p>\n<p>What there is to diminish is the reading habit that dominates clinical medicine: accepting AUROC as a verdict of quality for outcomes that are, almost all of them, rare. The question that matters is not &#8220;does this model have a high AUC?&#8221; but <em>&#8220;what is the prevalence, and what happens to the team when that alert fires for the third time in the same night?&#8221;<\/em>.<\/p>\n<p>\ud83d\udca1 <strong>Connecting the Dots:<\/strong> the authority asset here is not the Average Precision of 0.60 \u2014 it is the fact that <strong>AUROC is mathematically insensitive to prevalence<\/strong>, and therefore the same 0.95 describes a model with 90% PPV and another with 10% PPV. Everyone cites the AUROC; almost no one in Brazilian medicine translates that <em>it does not know whether the outcome is rare, and therefore knows nothing about what will happen on shift<\/em>. But the second leap is the one almost nobody makes: if high sensitivity with low PPV is the inevitable signature of a rare outcome, then the design variable is not the algorithm \u2014 it is <strong>the clinical act the output requests<\/strong>. &#8220;High risk&#8221; demands a PPV that a rare outcome cannot deliver; &#8220;reassessment within 15 minutes&#8221; does not. And because CFM 2.454 classifies by <em>degree of autonomy<\/em> rather than accuracy, rewriting that sentence on the screen simultaneously lowers the statistical requirement and the regulatory risk class. Clinical engineering and regulatory engineering turn out to be the same decision \u2014 and that decision is not taught in a machine learning course, because it requires knowing what happens to the nursing staff in month three, nor in residency, because it requires knowing why the precision-recall curve exists. It is exactly that intersection, and not access to the model, that constitutes the technical moat for anyone building clinical AI in Brazil.<\/p>\n<\/div>\n<\/section>\n<section class=\"section\" id=\"ref-en\">\n<div class=\"reference-box\">\n<h4>References<\/h4>\n<ol>\n<li>Ha T, Kappy B, Chamberlain JM, McKinley KW. <em>Early Prediction of Critical Care Interventions From Pediatric Emergency Department Triage<\/em>. Hosp Pediatr. 2026 Aug;16(8):e618\u2013e625. (886,183 visits; 26,721 outcomes, 3.0%; neural network AP 0.60, 95% CI 0.59\u20130.61; sensitivity 88%, PPV 32%; counterfactual in Table 4) Available at: https:\/\/doi.org\/10.1542\/hpeds.2025-009127<\/li>\n<li>Davis J, Goadrich M. <em>The Relationship Between Precision-Recall and ROC Curves<\/em>. Proceedings of the 23rd International Conference on Machine Learning (ICML). 2006:233\u2013240. Available at: https:\/\/doi.org\/10.1145\/1143844.1143874<\/li>\n<li>Saito T, Rehmsmeier M. <em>The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets<\/em>. PLoS One. 2015;10(3):e0118432. (open access; see p. 5 for the PRC baseline, y = P\/(P+N), and Fig 2B for the visual demonstration) Available at: https:\/\/doi.org\/10.1371\/journal.pone.0118432<\/li>\n<li>Federal Council of Medicine (CFM). <em>Resolution CFM No. 2,454, of February 11, 2026 \u2014 Regulating the use of artificial intelligence in medicine<\/em>. Official Gazette 2026 Feb 27; Ed. 39, Sec. 1. Available at: https:\/\/sistemas.cfm.org.br\/normas\/arquivos\/resolucoes\/BR\/2026\/2454_2026.pdf<\/li>\n<li>Jiang X, Yang N, Shi T, et al. <em>Deciphering the &#8220;non-verbal&#8221; code: A preliminary exploration of multimodal large language models for neonatal pain recognition<\/em>. Digit Health. 2026;12:20552076261473729. (same sensitivity\/specificity asymmetry in a distinct domain) Available at: https:\/\/doi.org\/10.1177\/20552076261473729<\/li>\n<li>Lonsdale H, Patel K, Domenico H, et al. <em>Development and external validation of the NEO-READY model to predict date of discharge among premature neonatal intensive care patients<\/em>. J Perinatol. 2026. Available at: https:\/\/doi.org\/10.1038\/s41372-026-02827-2<\/li>\n<li><em>Artificial Intelligence in Pediatric Cardiac Intensive Care: Clinical Applications, Implementation Challenges, and Future Directions<\/em>. Curr Treat Options Pediatr. 2026. Available at: https:\/\/link.springer.com\/article\/10.1007\/s40746-026-00374-8<\/li>\n<\/ol>\n<\/div>\n<\/section>\n<\/div>\n<\/div>\n<p><!-- FOOTER --><\/p>\n<div class=\"footer\">\n<p><strong>Dr. Mbula Luzingu Barros<\/strong><\/p>\n<p id=\"footer-specialty-pt\">M\u00e9dico Pediatra Intensivista \u00b7 25 anos de UTI | Consultor em IA na Sa\u00fade | Fundador da INOVAMED | Criador da Metodologia AIMED<\/p>\n<p id=\"footer-specialty-en\" class=\"hidden\">Pediatric Intensivist \u00b7 25 years in ICU | AI Healthcare Consultant | Founder of INOVAMED | Creator of the AIMED Methodology<\/p>\n<p>\u00a9 2026 inovamed.pro \u2014 Este artigo pode ser compartilhado com atribui\u00e7\u00e3o ao autor \u00b7 This article may be shared with attribution to the author<\/p>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Um modelo treinado em 886 mil atendimentos pedi\u00e1tricos identifica, s\u00f3 com dados da triagem, 88% das crian\u00e7as que v\u00e3o precisar de suporte avan\u00e7ado. Mas a m\u00e9trica que consagrou esses modelos \u00e9 cega para a raridade do evento \u2014 e entender por qu\u00ea separa quem consome escore preditivo de quem sabe julg\u00e1-lo.<span class=\"more-link\"><a href=\"https:\/\/inovamed.pro\/?p=2805\">LEIA O ARTIGO COMPLETO<\/a><\/span><\/p>\n","protected":false},"author":1,"featured_media":2808,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[1],"tags":[],"class_list":["entry","author-mbulabarros","has-excerpt","post-2805","post","type-post","status-publish","format-standard","has-post-thumbnail","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v23.9 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9 - INOVAMED<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/inovamed.pro\/?p=2805\" \/>\n<meta property=\"og:locale\" content=\"pt_BR\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9 - INOVAMED\" \/>\n<meta property=\"og:description\" content=\"Um modelo treinado em 886 mil atendimentos pedi\u00e1tricos identifica, s\u00f3 com dados da triagem, 88% das crian\u00e7as que v\u00e3o precisar de suporte avan\u00e7ado. Mas a m\u00e9trica que consagrou esses modelos \u00e9 cega para a raridade do evento \u2014 e entender por qu\u00ea separa quem consome escore preditivo de quem sabe julg\u00e1-lo.LEIA O ARTIGO COMPLETO\" \/>\n<meta property=\"og:url\" content=\"https:\/\/inovamed.pro\/?p=2805\" \/>\n<meta property=\"og:site_name\" content=\"INOVAMED\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-28T21:09:27+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-28T21:36:16+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"630\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"mbulabarros\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Escrito por\" \/>\n\t<meta name=\"twitter:data1\" content=\"mbulabarros\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. tempo de leitura\" \/>\n\t<meta name=\"twitter:data2\" content=\"54 minutos\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/inovamed.pro\/?p=2805#article\",\"isPartOf\":{\"@id\":\"https:\/\/inovamed.pro\/?p=2805\"},\"author\":{\"name\":\"mbulabarros\",\"@id\":\"https:\/\/inovamed.pro\/#\/schema\/person\/f023d988086a844cf27e25d0dd97e239\"},\"headline\":\"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9\",\"datePublished\":\"2026-07-28T21:09:27+00:00\",\"dateModified\":\"2026-07-28T21:36:16+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/inovamed.pro\/?p=2805\"},\"wordCount\":10941,\"publisher\":{\"@id\":\"https:\/\/inovamed.pro\/#organization\"},\"image\":{\"@id\":\"https:\/\/inovamed.pro\/?p=2805#primaryimage\"},\"thumbnailUrl\":\"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg\",\"articleSection\":[\"Uncategorized\"],\"inLanguage\":\"pt-BR\"},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/inovamed.pro\/?p=2805\",\"url\":\"https:\/\/inovamed.pro\/?p=2805\",\"name\":\"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9 - INOVAMED\",\"isPartOf\":{\"@id\":\"https:\/\/inovamed.pro\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/inovamed.pro\/?p=2805#primaryimage\"},\"image\":{\"@id\":\"https:\/\/inovamed.pro\/?p=2805#primaryimage\"},\"thumbnailUrl\":\"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg\",\"datePublished\":\"2026-07-28T21:09:27+00:00\",\"dateModified\":\"2026-07-28T21:36:16+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/inovamed.pro\/?p=2805#breadcrumb\"},\"inLanguage\":\"pt-BR\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/inovamed.pro\/?p=2805\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"pt-BR\",\"@id\":\"https:\/\/inovamed.pro\/?p=2805#primaryimage\",\"url\":\"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg\",\"contentUrl\":\"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg\",\"width\":1200,\"height\":630,\"caption\":\"Ilustra\u00e7\u00e3o editorial de um pronto-socorro pedi\u00e1trico vazio \u00e0 noite, em tons de azul e verde-azulado\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/inovamed.pro\/?p=2805#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/inovamed.pro\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/inovamed.pro\/#website\",\"url\":\"https:\/\/inovamed.pro\/\",\"name\":\"INOVAMED\",\"description\":\"Intelig\u00eancia Artificial em Medicina\",\"publisher\":{\"@id\":\"https:\/\/inovamed.pro\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/inovamed.pro\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"pt-BR\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/inovamed.pro\/#organization\",\"name\":\"INOVAMED\",\"url\":\"https:\/\/inovamed.pro\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"pt-BR\",\"@id\":\"https:\/\/inovamed.pro\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/inovamed.pro\/wp-content\/uploads\/2024\/12\/cropped-DALL\u00b7E-2024-12-10-13.53.06-A-hyperrealistic-3D-rendering-of-a-half-human-brain-merged-with-circuit-board-patterns-viewed-from-a-slight-side-angle.-The-brains-left-hemisphere-s.webp\",\"contentUrl\":\"https:\/\/inovamed.pro\/wp-content\/uploads\/2024\/12\/cropped-DALL\u00b7E-2024-12-10-13.53.06-A-hyperrealistic-3D-rendering-of-a-half-human-brain-merged-with-circuit-board-patterns-viewed-from-a-slight-side-angle.-The-brains-left-hemisphere-s.webp\",\"width\":1706,\"height\":1023,\"caption\":\"INOVAMED\"},\"image\":{\"@id\":\"https:\/\/inovamed.pro\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/inovamed.pro\/#\/schema\/person\/f023d988086a844cf27e25d0dd97e239\",\"name\":\"mbulabarros\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"pt-BR\",\"@id\":\"https:\/\/inovamed.pro\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/b45ea8fc03f19632f2c0e1180c867958fb387b48b3fe9af6e41ad80bbcf0d3c9?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/b45ea8fc03f19632f2c0e1180c867958fb387b48b3fe9af6e41ad80bbcf0d3c9?s=96&d=mm&r=g\",\"caption\":\"mbulabarros\"},\"sameAs\":[\"https:\/\/inovamed.pro\"],\"url\":\"https:\/\/inovamed.pro\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9 - INOVAMED","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/inovamed.pro\/?p=2805","og_locale":"pt_BR","og_type":"article","og_title":"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9 - INOVAMED","og_description":"Um modelo treinado em 886 mil atendimentos pedi\u00e1tricos identifica, s\u00f3 com dados da triagem, 88% das crian\u00e7as que v\u00e3o precisar de suporte avan\u00e7ado. Mas a m\u00e9trica que consagrou esses modelos \u00e9 cega para a raridade do evento \u2014 e entender por qu\u00ea separa quem consome escore preditivo de quem sabe julg\u00e1-lo.LEIA O ARTIGO COMPLETO","og_url":"https:\/\/inovamed.pro\/?p=2805","og_site_name":"INOVAMED","article_published_time":"2026-07-28T21:09:27+00:00","article_modified_time":"2026-07-28T21:36:16+00:00","og_image":[{"width":1200,"height":630,"url":"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg","type":"image\/jpeg"}],"author":"mbulabarros","twitter_card":"summary_large_image","twitter_misc":{"Escrito por":"mbulabarros","Est. tempo de leitura":"54 minutos"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/inovamed.pro\/?p=2805#article","isPartOf":{"@id":"https:\/\/inovamed.pro\/?p=2805"},"author":{"name":"mbulabarros","@id":"https:\/\/inovamed.pro\/#\/schema\/person\/f023d988086a844cf27e25d0dd97e239"},"headline":"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9","datePublished":"2026-07-28T21:09:27+00:00","dateModified":"2026-07-28T21:36:16+00:00","mainEntityOfPage":{"@id":"https:\/\/inovamed.pro\/?p=2805"},"wordCount":10941,"publisher":{"@id":"https:\/\/inovamed.pro\/#organization"},"image":{"@id":"https:\/\/inovamed.pro\/?p=2805#primaryimage"},"thumbnailUrl":"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg","articleSection":["Uncategorized"],"inLanguage":"pt-BR"},{"@type":"WebPage","@id":"https:\/\/inovamed.pro\/?p=2805","url":"https:\/\/inovamed.pro\/?p=2805","name":"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9 - INOVAMED","isPartOf":{"@id":"https:\/\/inovamed.pro\/#website"},"primaryImageOfPage":{"@id":"https:\/\/inovamed.pro\/?p=2805#primaryimage"},"image":{"@id":"https:\/\/inovamed.pro\/?p=2805#primaryimage"},"thumbnailUrl":"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg","datePublished":"2026-07-28T21:09:27+00:00","dateModified":"2026-07-28T21:36:16+00:00","breadcrumb":{"@id":"https:\/\/inovamed.pro\/?p=2805#breadcrumb"},"inLanguage":"pt-BR","potentialAction":[{"@type":"ReadAction","target":["https:\/\/inovamed.pro\/?p=2805"]}]},{"@type":"ImageObject","inLanguage":"pt-BR","@id":"https:\/\/inovamed.pro\/?p=2805#primaryimage","url":"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg","contentUrl":"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg","width":1200,"height":630,"caption":"Ilustra\u00e7\u00e3o editorial de um pronto-socorro pedi\u00e1trico vazio \u00e0 noite, em tons de azul e verde-azulado"},{"@type":"BreadcrumbList","@id":"https:\/\/inovamed.pro\/?p=2805#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/inovamed.pro\/"},{"@type":"ListItem","position":2,"name":"A IA Que Acha a Crian\u00e7a Grave na Fila do Pronto-Socorro \u2014 e a M\u00e9trica Que Faz Todo Modelo Parecer Melhor do Que \u00c9"}]},{"@type":"WebSite","@id":"https:\/\/inovamed.pro\/#website","url":"https:\/\/inovamed.pro\/","name":"INOVAMED","description":"Intelig\u00eancia Artificial em Medicina","publisher":{"@id":"https:\/\/inovamed.pro\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/inovamed.pro\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"pt-BR"},{"@type":"Organization","@id":"https:\/\/inovamed.pro\/#organization","name":"INOVAMED","url":"https:\/\/inovamed.pro\/","logo":{"@type":"ImageObject","inLanguage":"pt-BR","@id":"https:\/\/inovamed.pro\/#\/schema\/logo\/image\/","url":"https:\/\/inovamed.pro\/wp-content\/uploads\/2024\/12\/cropped-DALL\u00b7E-2024-12-10-13.53.06-A-hyperrealistic-3D-rendering-of-a-half-human-brain-merged-with-circuit-board-patterns-viewed-from-a-slight-side-angle.-The-brains-left-hemisphere-s.webp","contentUrl":"https:\/\/inovamed.pro\/wp-content\/uploads\/2024\/12\/cropped-DALL\u00b7E-2024-12-10-13.53.06-A-hyperrealistic-3D-rendering-of-a-half-human-brain-merged-with-circuit-board-patterns-viewed-from-a-slight-side-angle.-The-brains-left-hemisphere-s.webp","width":1706,"height":1023,"caption":"INOVAMED"},"image":{"@id":"https:\/\/inovamed.pro\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/inovamed.pro\/#\/schema\/person\/f023d988086a844cf27e25d0dd97e239","name":"mbulabarros","image":{"@type":"ImageObject","inLanguage":"pt-BR","@id":"https:\/\/inovamed.pro\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/b45ea8fc03f19632f2c0e1180c867958fb387b48b3fe9af6e41ad80bbcf0d3c9?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/b45ea8fc03f19632f2c0e1180c867958fb387b48b3fe9af6e41ad80bbcf0d3c9?s=96&d=mm&r=g","caption":"mbulabarros"},"sameAs":["https:\/\/inovamed.pro"],"url":"https:\/\/inovamed.pro\/?author=1"}]}},"jetpack_featured_media_url":"https:\/\/inovamed.pro\/wp-content\/uploads\/2026\/07\/triagem-pediatrica-ia-metrica-1200.jpg","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/inovamed.pro\/index.php?rest_route=\/wp\/v2\/posts\/2805","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/inovamed.pro\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/inovamed.pro\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/inovamed.pro\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/inovamed.pro\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2805"}],"version-history":[{"count":1,"href":"https:\/\/inovamed.pro\/index.php?rest_route=\/wp\/v2\/posts\/2805\/revisions"}],"predecessor-version":[{"id":2806,"href":"https:\/\/inovamed.pro\/index.php?rest_route=\/wp\/v2\/posts\/2805\/revisions\/2806"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/inovamed.pro\/index.php?rest_route=\/wp\/v2\/media\/2808"}],"wp:attachment":[{"href":"https:\/\/inovamed.pro\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2805"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/inovamed.pro\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2805"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/inovamed.pro\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2805"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}