Metade do Tempo, Melhor Desempenho — e o Efeito Grande Demais Para Ser Verdade
Um ensaio randomizado com 124 residentes mostrou que treinar com inteligência artificial reduziu o estudo semanal de 20,6 para 10,6 horas e melhorou o desempenho em todos os domínios. É o número que rende manchete. Mas há um dado que quase nenhum texto traduz: o estudo equivalente em ginecologia e obstetrícia reporta um tamanho de efeito de d = 2,30 — cinco vezes o típico da pesquisa educacional — em um desenho sem grupo controle. Entender por que um efeito grande demais é um sinal de alerta, e não de sucesso, é o que separa quem consome evidência educacional de quem sabe julgá-la.
📅 Publicado em 11 de agosto de 2026
Navegue pelo Artigo
Do ensaio randomizado ao efeito grande demais — por que julgar evidência educacional é uma competência clínica
- 1. Por Que Isso Importa Para Quem Ensina Residente
- 2. O Que os Três Estudos Entregaram
- 3. Os Números, Sem Filtro
- 4. A Armadilha: O Efeito Grande Demais
- 5. Instrumentar É um Ato Educacional
- 6. A Camada Regulatória — o Vazio Que Ninguém Nomeia
- 7. A Linha do Tempo
- 8. O Que Fazer na Segunda-Feira
- 9. Exercício: E Se Fosse Cardiologia?
- 10. Considerações Finais
- 11. Referências
Para quem não é da área — o essencial em 60 segundos
Este artigo funciona em duas camadas. A técnica, para quem coordena programa de residência ou pesquisa educação médica; e a geral, que só exige curiosidade. Toda ideia central aparece duas vezes: uma com exemplo de hospital, outra com exemplo do dia a dia.
Residência médica é o treinamento em serviço que um médico faz depois de formado, de 2 a 6 anos, dentro de um hospital, sob supervisão. É onde a medicina de verdade é aprendida — e o recurso mais escasso ali não é informação, é tempo de supervisão qualificada.
Ensaio clínico randomizado (RCT) é o desenho em que os participantes são sorteados para receber ou não a intervenção. O sorteio é o que permite dizer que a diferença observada foi causada pela intervenção, e não por diferenças prévias entre os grupos.
Estudo pré-pós mede o mesmo grupo antes e depois, sem grupo de comparação. É mais fácil de fazer e muito mais fraco: qualquer coisa que acontecesse naquele período — inclusive simplesmente amadurecer — aparece misturada ao efeito da intervenção.
Tamanho de efeito (Cohen’s d) é o quanto os grupos se afastam, medido em desvios-padrão. Por convenção, 0,2 é pequeno, 0,5 é médio, 0,8 é grande. É a palavra mais importante deste artigo — guarde essa.
Por Que Isso Importa Para Quem Ensina Residente
Todo programa de residência vive a mesma escassez, e ela não é de conteúdo. Diretriz é pública, artigo é acessível, videoaula é gratuita. O que falta é tempo protegido de raciocínio supervisionado — o preceptor sentado ao lado do residente, olhando o mesmo caso, corrigindo o mesmo erro na hora em que ele acontece.
Esse recurso não escala. Um preceptor consegue dar feedback denso e individualizado a três, talvez quatro residentes por rotação. Acima disso, o feedback vira genérico, atrasado, ou simplesmente não acontece. É por isso que a educação médica competency-based, adotada como padrão internacional pela ACGME e pelo Royal College canadense, tem uma barreira de implementação declarada na própria literatura: tempo docente indisponível para mentoria e avaliação individualizadas.
Ministrei duas turmas do AIMED para residentes de ginecologia e obstetrícia em maio e junho deste ano, no hospital onde esses residentes fazem sua formação. Foram turmas mistas, do primeiro ao terceiro ano de residência na mesma sala. Foi essa experiência que me levou a ler os ensaios publicados em 2026 sobre exatamente isso. E o que encontrei não foi o que eu esperava — nem no sentido de decepção, nem no de confirmação.
O Que os Três Estudos Entregaram
Estudo 1 — o ensaio randomizado (oncologia mamária, n = 124)
Publicado em 12 de março de 2026 na Frontiers in Medicine. Cento e vinte e quatro residentes de oncologia mamária de três hospitais terciários urbanos da mesma região metropolitana, todos centros credenciados para treinamento padronizado. Randomização individual 1:1 por sequência gerada em computador, preparada por estatístico independente, com ocultação de alocação por envelopes opacos numerados sequencialmente. Sessenta e dois no braço AIEIT, sessenta e dois em treinamento convencional.
Os avaliadores dos desfechos subjetivos — raciocínio clínico, apresentação de caso, trabalho em equipe — foram cegados para a alocação e não participaram da entrega da intervenção. Participantes e instrutores não foram cegados, o que é inevitável em intervenção educacional. Cálculo amostral prévio: efeito esperado d ≈ 0,6, alfa bicaudal 0,05, poder 80%, exigindo 58 por grupo. A intervenção durou uma rotação clínica de três meses em tempo integral.
A intervenção tinha quatro componentes: grafo de conhecimento dinâmico com processamento de linguagem natural sobre diretrizes NCCN, atualizado mensalmente; plataforma de reunião multidisciplinar em realidade mista, com cerca de 20 cenários padronizados; sistema de paciente virtual com mentor de IA, com cerca de 25 casos; e um painel de learning analytics. Desfecho primário declarado: aquisição de conhecimento e desempenho em raciocínio clínico.
Estudo 2 — a simulação de ecocardiografia transesofágica (anestesiologia, n = 60)
Publicado em 25 de maio de 2026 na Frontiers in Surgery. Ensaio randomizado de centro único, no Hospital Provincial do Povo de Sichuan, entre dezembro de 2024 e dezembro de 2025. Sessenta residentes de anestesiologia, trinta em cada braço. Os dois grupos receberam o mesmo currículo de ETE conduzido por instrutor — o braço de IA recebeu, adicionalmente, treino em simulador com anatomia 3D, exercícios interativos de leitura de imagem e manipulação virtual da sonda com feedback em tempo real.
Esse desenho é o mais limpo dos três, e vale entender por quê: como a base curricular é idêntica, a diferença observada é atribuível ao acréscimo do simulador, não a “ter tecnologia versus não ter nada”.
Estudo 3 — o de ginecologia e obstetrícia (n = 28)
Publicado em 2026 na BMC Medical Education. E aqui está a diferença que muda tudo: não é ensaio randomizado. É um estudo prospectivo intervencional com desenho pré-pós, sem grupo controle. Vinte e oito residentes de três coortes consecutivas — 2021 (n = 10), 2022 (n = 10), 2023 (n = 8) — de um único hospital, o Primeiro Hospital Afiliado da Universidade Médica de Fujian. O período total foi de 36 meses, e a comparação principal cobre 12 meses de treinamento na plataforma.
Mesma família de intervenção — plataforma de simulação de casos, trilha personalizada por algoritmo adaptativo, sistema de avaliação por competência. Um degrau inteiro abaixo na hierarquia de evidência.
Um exemplo concreto: o que “metade do tempo” significa aqui
No ensaio de oncologia, o grupo controle estudou 20,62 ± 2,48 horas por semana. O grupo de IA estudou 10,63 ± 2,54 horas — e foi melhor em todos os domínios avaliados.
Traduzindo para o chão da residência: são dez horas semanais devolvidas a cada residente. Ao longo de uma rotação de três meses, isso é aproximadamente 130 horas por pessoa. Num programa com vinte residentes, 2.600 horas. Esse é o ativo real — não o ganho de nota.
E é aqui que o número precisa de contexto brasileiro: a carga horária da residência médica no Brasil já é de 60 horas semanais, das quais 10% a 20% devem ser atividades teóricas. Vinte horas semanais de estudo adicional, como no braço controle chinês, não é o padrão da maior parte dos nossos programas — é acima dele. O denominador do ganho, portanto, não é o mesmo.
Os Números, Sem Filtro
A tabela abaixo separa o que foi medido em ensaio randomizado do que foi medido sem controle. Essa separação é o conteúdo principal do artigo — não os valores em si.
| Desfecho | Intervenção | Controle | Desenho · Fonte | Leitura |
|---|---|---|---|---|
| Tempo de estudo semanal (h) | 10,63 ± 2,54 | 20,62 ± 2,48 | RCT · Ji et al. · 2026 | ↑ metade do tempo |
| Classificação molecular (% acerto) | 93,66 ± 1,92 | 82,65 ± 2,10 | RCT · Ji et al. · 2026 | ↑ p < 0,001 |
| Biópsia guiada por US — tempo (min) | 8,5 ± 2,1 | 14,2 ± 3,8 | RCT · Ji et al. · 2026 | ↑ −40% |
| Retenção de conhecimento em 6 meses (%) | 86,7 ± 4,2 | 74,2 ± 3,2 | RCT · Ji et al. · 2026 | ↑ sustentado |
| Aprovação em prova de especialidade (%) | 90,3 | 67,6 | RCT · Ji et al. · 2026 | ↑ p = 0,0036 |
| Incidência de burnout (%) | 35,2 | 61,5 | RCT · Ji et al. · 2026 | ↑ subjetivo, não cegado |
| Complicações pós-operatórias, satisfação | favorável ao braço de IA | RCT · exploratório | ⚠ os autores pedem cautela | |
| ETE — aquisição de janelas-chave (%) | 88,2 ± 8,1 | 76,5 ± 9,7 | RCT · Zhang et al. · 2026 | ↑ base curricular idêntica |
| ETE — OSATS (0–150) | 121,5 ± 9,2 | 112,4 ± 10,1 | RCT · Zhang et al. · 2026 | ↑ p < 0,001 |
| G.O. — raciocínio clínico (pontos) | 68,4 ± 8,1 → 84,9 ± 5,7 · d = 2,36 | pré-pós, sem controle · 2026 | ⚠ ver seção 4 | |
| G.O. — competência global (pontos) | 72,9 ± 6,8 → 86,4 ± 4,8 · d = 2,30 | pré-pós, sem controle · 2026 | ⚠ ver seção 4 | |
O tradutor de tamanho de efeito — comece por aqui
Cohen’s d mede a distância entre duas médias em unidades de desvio-padrão. A convenção de Cohen: 0,2 pequeno, 0,5 médio, 0,8 grande.
No hospital: se você troca um antitérmico por outro e a temperatura média cai meio grau a mais, com um desvio-padrão de um grau, isso é d = 0,5. Perceptível, mas não transformador.
Fora dele: a diferença de altura média entre homens e mulheres adultos é de aproximadamente d = 2,0. É uma diferença que você enxerga de longe, sem medir ninguém.
Um d = 2,30 em intervenção educacional significa que o residente médio depois do treino supera cerca de 99% dos residentes antes do treino. Na síntese de Hattie sobre mais de oitocentas meta-análises em educação, o efeito médio de qualquer intervenção pedagógica fica em torno de d = 0,40. Um efeito cinco a seis vezes maior que a média da literatura não é motivo para comemorar — é motivo para procurar a explicação alternativa.
A Armadilha: O Efeito Grande Demais
Todos os quatro domínios do estudo de G.O. reportam Cohen’s d entre 1,58 e 2,56. Competência global: 2,30. Habilidades clínicas: 2,56. Raciocínio clínico: 2,36. Conhecimento profissional: 2,02.
Esses são efeitos gigantescos. E o desenho que os produziu é pré-pós sem grupo controle, medindo 12 meses de residência.
Pergunte-se o óbvio: o que mais aconteceu com esses 28 residentes ao longo de doze meses? Eles fizeram residência. Atenderam pacientes, passaram por plantões, foram supervisionados, leram, erraram e corrigiram. Um residente de primeiro ano melhora substancialmente do mês 1 ao mês 12 com ou sem plataforma de IA. Sem grupo controle, não existe nenhuma forma de separar o efeito da intervenção do efeito de simplesmente ter feito mais um ano de residência.
Some a isso dois mecanismos bem descritos que inflam sistematicamente o pré-pós: o efeito de testagem — quem já fez a avaliação uma vez vai melhor na segunda, mesmo sem aprender nada — e o efeito Hawthorne, a melhora de desempenho que decorre de saber que se está sendo observado dentro de um projeto de inovação. Os autores do ensaio randomizado de oncologia declaram explicitamente que não podem excluir o efeito Hawthorne no próprio ensaio deles, que tem grupo controle. No estudo sem controle, não há nem como levantar a hipótese.
Nada disso significa que a intervenção não funcione. Significa que o estudo de G.O. não consegue medir o quanto ela funciona — e que o d = 2,30 deve ser lido como o efeito combinado de intervenção, maturação, testagem e observação, não como o efeito da plataforma.
⚠ A ressalva que vale igualmente para o ensaio randomizado
Os autores do estudo de oncologia são exemplares em declarar os limites, e vale reproduzi-los: o estudo foi conduzido apenas em hospitais terciários urbanos, limitando a generalização para contextos de recurso limitado; a randomização não foi estratificada por centro, de modo que efeitos residuais de centro não podem ser excluídos; múltiplos desfechos foram avaliados sem ajuste formal para comparações múltiplas, aumentando o risco de erro tipo I; parte das avaliações é subjetiva; e os desfechos de paciente — complicações pós-operatórias, tempo de internação, satisfação — são declaradamente exploratórios e devem ser interpretados com cautela.
Há ainda uma limitação estrutural que os próprios autores nomeiam e que quase nenhum leitor vai reter: a intervenção é multimodal. Grafo de conhecimento, realidade mista, paciente virtual e painel de analytics foram entregues como pacote. O efeito observado não pode ser atribuído à IA isoladamente — ele reflete a combinação. Os autores pedem, corretamente, estudos fatoriais futuros para isolar cada componente.
⚠ O bloqueio de transferência para o Brasil que ninguém vai notar
Os três estudos são chineses. Todos os três. Oncologia em Guangzhou, anestesiologia em Sichuan, ginecologia e obstetrícia em Fujian.
Isso não é um problema de qualidade — são estudos bem conduzidos e aprovados por comitês de ética institucionais. É um problema de base de evidência de país único. A residência médica chinesa tem estrutura curricular, razão preceptor-residente, cultura de avaliação e infraestrutura tecnológica próprias. Nenhuma dessas variáveis é neutra em relação ao tamanho do efeito de uma intervenção educacional.
E há uma segunda camada, mais desconfortável: em todos os três casos, a plataforma avaliada foi construída pelo próprio grupo que a avaliou. No estudo de oncologia, os autores escrevem literalmente “nosso framework proposto”. Isso não é fraude nem má-fé — é a norma em pesquisa de inovação educacional, onde quem constrói é quem testa. Mas é exatamente a configuração em que a literatura de eficácia costuma superestimar efeitos, e é a razão pela qual replicação independente é o próximo passo obrigatório, não um refinamento opcional.
Não existe, até onde consegui verificar, um único estudo brasileiro publicado com esse desenho. Zero RCT, zero pré-pós indexado, em qualquer especialidade.
Instrumentar É um Ato Educacional
Aqui a leitura vira para o lado prático, e a conclusão é contraintuitiva.
Olhe o que a intervenção do ensaio randomizado tem de caro: realidade mista, vinte cenários de reunião multidisciplinar, vinte e cinco pacientes virtuais, processamento de linguagem natural sobre diretrizes atualizado mensalmente. É um projeto de engenharia de meses, com orçamento.
Agora olhe o que os próprios autores identificam como mecanismo do efeito, na discussão: personalização do caminho de aprendizado, feedback imediato em ambiente sem risco, e a passagem de avaliação somativa episódica para monitoramento contínuo com remediação dirigida. Traduzindo: feedback denso, individualizado, de alta frequência.
Isso é exatamente o que um bom preceptor faz. O que a plataforma resolve não é a natureza do feedback — é a escala dele. E o que a IA barateou de forma mais radical não foi o ensino: foi a instrumentação do ensino. Medir competência antes e depois, com rubrica validada e análise pareada, custa hoje uma planilha e trinta linhas de código.
Aplique isso ao que acontece nas suas próprias turmas. Duas turmas do AIMED para residentes de G.O., em maio e junho, mistas do primeiro ao terceiro ano — se tivessem sido instrumentadas com uma rubrica aplicada na entrada e na saída, teriam gerado o mesmo nível de evidência do estudo de Fujian, com duas vantagens que ele não tem: duas coortes independentes em vez de três coortes do mesmo serviço, e — a mais importante — o ano de residência como variável de estratificação limpa em vez de confundidor.
A diferença entre dado e publicação, nesse caso, não é rigor metodológico. É instrumentação — e instrumentação é uma decisão que se toma antes da aula começar, não depois.
💡 A turma mista resolve um confundimento que o estudo publicado não conseguiu resolver
O achado mais elegante do estudo de Fujian é a convergência entre coortes. Antes do treinamento, as três coortes tinham escores basais significativamente diferentes (ANOVA, F = 4,25; p = 0,025) — o que é esperado, já que correspondem a terceiro, segundo e primeiro ano. Depois, a diferença desapareceu (F = 2,14; p = 0,138), com redução de 47,5% na variância entre coortes (DP 4,04 → 2,12). Os autores interpretam isso como evidência de que o modelo adaptativo atende necessidades distintas de estágios distintos, padronizando a competência final.
A interpretação é plausível. Mas o desenho não consegue sustentá-la, e vale entender exatamente por quê: naquele estudo, ano de residência está inteiramente confundido com período de calendário. A coorte de 2021 não é apenas “o terceiro ano” — é um grupo de pessoas diferentes, recrutado em outro momento, com exposição prévia distinta e treinado num período diferente. Não há como separar “o treinamento comprimiu a variância entre níveis” de “estes três grupos de pessoas simplesmente convergiram”.
Uma turma mista de R1 a R3 na mesma sala elimina esse confundimento por construção. Mesmo instrutor, mesmo dia, mesmo conteúdo, mesma sequência de casos. O ano de residência deixa de ser um confundidor e passa a ser um fator de estratificação limpo — e a pergunta “o ganho é maior em quem começa mais atrás?” fica respondível com uma ANOVA de dois fatores banal.
Some as duas turmas de maio e junho e você tem replicação independente de um efeito de estratificação. Isso não é equivalente ao que está publicado na sua especialidade. É metodologicamente superior — em um ponto específico, mas justamente naquele que os autores de Fujian elegeram como o achado mais interessante do trabalho deles.
Um exemplo concreto: o desenho em cunha, ou como ter controle sem negar treinamento
A objeção ética óbvia ao grupo controle em educação é: “não vou negar treinamento a metade da minha turma”. Correto — e desnecessário.
No desenho em cunha (stepped wedge), todos os participantes recebem a intervenção, mas em momentos diferentes. Divida a turma em três subgrupos e entregue os módulos em ordem escalonada: o subgrupo A recebe o módulo 1 no mês 1 e o módulo 2 no mês 2; o subgrupo B recebe na ordem inversa; e assim por diante.
Cada subgrupo funciona como controle dos outros no período em que ainda não recebeu aquele módulo. Ninguém fica sem nada. E você ganha um comparador interno que o estudo de Fujian não tem — o que colocaria a sua coorte, metodologicamente, acima do que está publicado na sua especialidade.
A Camada Regulatória — o Vazio Que Ninguém Nomeia
A Resolução CFM nº 2.454/2026, publicada no Diário Oficial da União em 27 de fevereiro de 2026, normatiza o uso da inteligência artificial na medicina. Ela classifica sistemas por nível de risco — baixo, médio, alto ou inaceitável —, garante que a palavra final sobre diagnóstico, terapêutica e prognóstico é sempre do médico, proíbe delegar à IA a comunicação de diagnósticos e prognósticos, e veda que instituições penalizem o médico que não seguir a recomendação da máquina.
Note o que ela regula: decisão clínica. E note o que ela, por construção, não alcança: IA aplicada à formação médica. Um simulador de paciente virtual que treina residente não emite diagnóstico sobre paciente real. Um painel de learning analytics que rastreia competência não influencia conduta terapêutica. Nenhum dos três estudos deste artigo descreve um sistema que a CFM 2.454 classificaria.
Na ANVISA, o raciocínio é o mesmo por outro caminho. A regularização de software como dispositivo médico depende de o software auxiliar diagnóstico, monitorar condição clínica ou influenciar decisão terapêutica. Plataforma educacional que opera sobre caso simulado não atende a esse gatilho.
A leitura correta disso não é “está liberado, façam o que quiserem”. É o oposto: é a área de IA médica com maior velocidade de adoção e menor densidade regulatória simultaneamente. Quando um programa de residência adota uma plataforma que decide o que cada residente estuda e como cada residente é avaliado, não há resolução do CFM nem registro da ANVISA que exija validação prévia. A única barreira é a qualidade da evidência que o programa exigir antes de adotar — e essa barreira é institucional, voluntária, e hoje praticamente inexistente.
⚠ Onde a fronteira volta a ser regulada
Há um ponto em que o argumento acima deixa de valer, e vale marcá-lo: no momento em que a plataforma educacional passa a operar sobre dados de paciente real — prontuário, imagem de exame do próprio serviço, desfecho clínico atribuído a um residente específico —, ela sai do domínio puramente educacional. Aí entram, conforme o caso, a LGPD, a exigência de aprovação em Comitê de Ética em Pesquisa e, dependendo da função, a discussão de SaMD.
O ensaio de oncologia lidou com isso da forma correta: os algoritmos adaptativos foram treinados exclusivamente sobre dados agregados e anonimizados de interação aluno-conteúdo, sem acesso a dados de paciente ou a prontuário eletrônico, e o sistema não gerava recomendação diagnóstica ou terapêutica autônoma. Essa fronteira é desenho de projeto, não detalhe de implementação.
A Linha do Tempo
2021–2024 · A fase pré-evidência
IA na educação médica aparece como ferramenta isolada e fragmentada. Simulação de realidade virtual, paciente virtual e avaliação automatizada existem, mas sem desenho comparativo. A literatura é de viabilidade e aceitação, não de eficácia.
2026 · O primeiro corpo de evidência randomizada — e onde você está agora
Três estudos, três especialidades, um só país. Dois com randomização, um sem controle. Nenhuma replicação independente, nenhum dado fora da China, nenhum estudo brasileiro. Os efeitos são grandes; a incerteza sobre o quanto deles é real, também.
Próximos 12–24 meses · A janela de replicação
É o intervalo em que estudos fatoriais vão começar a isolar qual componente carrega o efeito, e em que replicações fora da China vão testar se o tamanho se sustenta. É também a janela em que uma coorte brasileira bem instrumentada entra na literatura como primeira, não como confirmação tardia.
Depois · A consolidação curricular
Quando a evidência estabilizar, a discussão migra de “funciona?” para “quanto da carga horária teórica pode ser redesenhada?”. Aí a decisão deixa de ser pedagógica e passa a ser de política de formação — CNRM, comissões de residência, colegiados.
O Que Fazer na Segunda-Feira
Se você coordena ou ensina em programa de residência
- ☐ Antes de adotar qualquer plataforma, pergunte pelo desenho do estudo que a sustenta. Se a resposta for “melhoramos a competência em X%”, pergunte: comparado com quem? Ausência de grupo controle é a informação mais importante do material de vendas — e é sempre a que não está nele.
- ☐ Desconfie de tamanho de efeito acima de d = 1,0 em intervenção educacional. O padrão da literatura é d ≈ 0,4. Efeito muito acima disso pede explicação alternativa antes de comemoração.
- ☐ Pergunte quem construiu a plataforma avaliada. Se quem publicou é quem desenvolveu, o achado continua válido — mas exige replicação independente antes de virar base de decisão institucional.
- ☐ Instrumente a próxima turma antes de ela começar. Rubrica de competência em quatro domínios, aplicada na entrada e na saída, com a mesma escala. Custo: uma planilha. Sem isso, o dado da turma se perde e não volta.
- ☐ Se quiser controle sem negar treinamento, use desenho em cunha. Ordem escalonada de módulos entre subgrupos entrega comparador interno com custo ético zero.
- ☐ Se a sua turma é mista por ano de residência, registre o ano de cada participante. É a variável mais barata de coletar e a de maior retorno analítico: com a entrega mantida constante, ela permite testar se o ganho é maior em quem começa mais atrás — pergunta que o estudo publicado na literatura não consegue responder sem confundimento.
- ☐ Registre tempo de estudo, não só desempenho. O achado mais robusto do ensaio randomizado é a redução de carga horária com desempenho preservado ou superior. É o desfecho de maior valor institucional e o mais fácil de medir.
- ☐ Submeta ao CEP antes de coletar. Dado educacional de participante identificável exige aprovação prévia para publicação. É o passo que mais frequentemente inviabiliza retrospectivamente uma coorte boa.
Exercício: E Se Fosse Cardiologia?
O padrão dos três estudos é transferível, mas não uniformemente. O que muda de especialidade para especialidade não é a tecnologia — é qual competência é o gargalo e qual é mensurável com rubrica objetiva. Vale fazer o exercício em voz alta, porque ele expõe o raciocínio de desenho melhor do que qualquer conclusão pronta.
Cardiologia: o caso mais favorável de todos, e quase ninguém percebeu
Cardiologia tem uma propriedade que oncologia e ginecologia não têm: três competências centrais com padrão-ouro objetivo, digitalizável e escalável.
Eletrocardiograma. É a única competência clínica de alto volume em que o item de avaliação é um arquivo, o gabarito é consensual, e o banco de casos pode ter dezenas de milhares de exemplos com desfecho conhecido. Não existe nada equivalente em ginecologia. Um sistema adaptativo que serve o próximo traçado calibrado pelo erro anterior do residente é, tecnicamente, a aplicação mais simples de todo este artigo — e a de maior densidade de repetição.
Ecocardiografia. Aqui já existe evidência direta e favorável, ainda que fora da cardiologia: o ensaio de ecocardiografia transesofágica com sessenta residentes de anestesiologia mostrou ganho em aquisição de janelas-chave (88,2% versus 76,5%) e redução de tempo de procedimento (27,1 versus 32,4 minutos), com base curricular idêntica nos dois braços. Ecocardiografia transtorácica em residência de cardiologia tem o mesmo perfil: aquisição de janela é medível objetivamente, e o OSATS é instrumento validado e transferível.
Hemodinâmica e eletrofisiologia. São os domínios em que a exposição real é mais escassa e mais dependente de volume institucional — exatamente a condição em que o simulador entrega mais valor marginal. E é onde a referência de competência acaba de ser atualizada: o Advanced Training Statement de eletrofisiologia cardíaca clínica de 2026, do ACC/AHA/HRS, revisou requisitos de treinamento e competências designadas com alinhamento explícito a tecnologias de ponta.
O detalhe que inverte a ordem de prioridade
A intuição diz “comece pelo ECG, é o mais fácil”. A leitura dos três estudos sugere o contrário.
Repare em qual domínio teve o maior ganho no estudo de G.O.: raciocínio clínico, que partiu do escore basal mais baixo (68,4) e subiu mais (16,5 pontos). Os autores registram que as habilidades cognitivas de ordem superior — acurácia de avaliação de risco e formulação de diagnóstico diferencial abrangente — tinham as menores notas iniciais e as maiores margens de melhora, 27,4% e 24,8%. O mesmo padrão aparece no ensaio de oncologia, cujo ganho mais expressivo foi em raciocínio oncológico especializado e colaboração multidisciplinar, e não em tarefa técnica isolada.
Ou seja: o efeito é maior onde a competência é mais complexa e menos treinável por repetição, não onde ela é mais fácil de digitalizar. Em cardiologia, isso desloca a aposta do ECG isolado para a decisão sob incerteza — a discussão de caso de síncope, a indicação de dispositivo, a estratificação pré-operatória, a conduta na dor torácica de risco intermediário. É o território onde a variabilidade entre cardiologistas é máxima e a rubrica é mais difícil de escrever. Também é onde o ganho, segundo os dados disponíveis, seria maior.
É um resultado desconfortável para quem planeja currículo, porque inverte a ordem de menor esforço. Mas é o que os dois estudos com dado por domínio mostram.
A pergunta para levar para o seu colegiado
Independentemente da especialidade, o exercício se reduz a três perguntas, nesta ordem:
1. Qual competência do meu programa tem o menor escore basal e a maior variabilidade entre residentes? (É onde o ganho será maior.)
2. Essa competência tem gabarito consensual entre os meus preceptores? (Se não tiver, o problema não é de plataforma — é de rubrica, e precisa ser resolvido antes.)
3. Consigo medi-la com a mesma régua na entrada e na saída? (Se não, não há estudo — há impressão.)
Se a resposta à pergunta 2 for não, você acabou de descobrir algo mais valioso que qualquer plataforma: a sua equipe não concorda sobre o que é o desempenho correto. Nenhum sistema de IA resolve isso, e todo sistema de IA vai amplificá-lo silenciosamente.
Aprenda a Julgar Evidência Antes de Adotar a Ferramenta
A metodologia AIMED forma médicos que constroem — não apenas consomem — ferramentas de IA clínica. Leitura crítica de desenho de estudo, tamanho de efeito sob desenho fraco, instrumentação de coorte própria e a fronteira regulatória entre educação e decisão clínica fazem parte do currículo, porque são a mesma competência vista de ângulos diferentes.
Considerações Finais
O ensaio de Ji e colaboradores é bom. Randomização por sequência computacional, ocultação de alocação por envelope opaco, avaliadores cegados para os desfechos subjetivos, cálculo amostral prévio e desfecho primário declarado. Reduzir o tempo de estudo semanal pela metade e ainda melhorar o desempenho é um achado que merece atenção, e os autores relatam as próprias limitações com honestidade acima da média.
O que merece crítica não são os estudos — é o hábito de leitura que vai se formar em torno deles. “IA melhora a formação do residente” vai circular como fato estabelecido, apoiado em três trabalhos de um país só, dois deles com o mesmo grupo avaliando a própria plataforma, um deles sem grupo controle e reportando um efeito cinco vezes maior que a média da literatura educacional.
💡 Connecting the Dots: o ativo de autoridade aqui não é o “metade do tempo” — é o fato de que o tamanho de efeito de uma intervenção educacional é inversamente proporcional à qualidade do desenho que o mediu, e o campo inteiro ainda lê efeito grande como boa notícia. O ensaio randomizado, com controle, relata diferenças expressivas mas plausíveis; o estudo sem controle relata d = 2,30, e é justamente o mais frágil que produz o número mais impressionante. Isso não é coincidência: é a assinatura previsível de um desenho que soma intervenção, maturação, testagem e efeito Hawthorne no mesmo número e chama tudo de efeito da plataforma. Mas o segundo salto é o que quase ninguém dá. Se o efeito é maior onde a competência é mais complexa e menos treinável por repetição — e os dois estudos com dado por domínio mostram exatamente isso —, então o alvo certo em qualquer especialidade não é a tarefa mais fácil de digitalizar; é a decisão sob incerteza, que é onde a rubrica é mais difícil de escrever porque os próprios preceptores não concordam entre si. Ou seja: a barreira para instrumentar a formação médica nunca foi tecnológica. É que a instrumentação obriga o serviço a declarar, por escrito e antes da aula, o que ele considera desempenho correto — e essa é uma conversa que a maioria dos programas nunca teve. Quem tiver essa conversa primeiro não ganha uma plataforma melhor: ganha a régua. E quem define a régua define o que a próxima geração da especialidade vai considerar competência.
Referências
- Ji F, Xiao W, Li X. AI-driven intelligent training enhances clinical competence in oncology residency: a randomized controlled trial. Front Med (Lausanne). 2026 Mar 12;13:1768388. (RCT, n = 124; tempo semanal 10,63 ± 2,54 h vs 20,62 ± 2,48 h; retenção 6 meses 86,7 ± 4,2% vs 74,2 ± 3,2%; aprovação em prova 90,3% vs 67,6%, p = 0,0036; burnout 35,2% vs 61,5%) Disponível em: https://doi.org/10.3389/fmed.2026.1768388
- Zhang M, Song SY, Zhou H, Xu G, Tong X, Fan D, Lei Q. Application of an artificial-intelligence-based transesophageal echocardiography simulation system in residency training. Front Surg. 2026 May 25;13:1762740. (RCT de centro único, n = 60; janelas-chave 88,2 ± 8,1% vs 76,5 ± 9,7%; OSATS-TEE 121,5 ± 9,2 vs 112,4 ± 10,1; tempo 27,1 ± 3,9 vs 32,4 ± 4,6 min) Disponível em: https://doi.org/10.3389/fsurg.2026.1762740
- Construction and implementation of an AI-enhanced progressive training model for clinical competency development in obstetrics and gynecology residents. BMC Med Educ. 2026. (estudo pré-pós sem grupo controle, n = 28, três coortes, centro único; competência global 72,9 ± 6,8 → 86,4 ± 4,8, d = 2,30; raciocínio clínico d = 2,36) Disponível em: https://doi.org/10.1186/s12909-026-08751-5
- Conselho Federal de Medicina. Resolução CFM nº 2.454, de 11 de fevereiro de 2026 — Normatiza o uso da inteligência artificial na medicina. DOU 2026 fev 27; Ed. 39, Seção 1. Disponível em: https://sistemas.cfm.org.br/normas/arquivos/resolucoes/BR/2026/2454_2026.pdf
- American College of Cardiology / American Heart Association / Heart Rhythm Society. 2026 ACC/AHA/HRS Advanced Training Statement on Clinical Cardiac Electrophysiology. J Am Coll Cardiol. 2026. (revisão dos requisitos de treinamento e competências designadas em eletrofisiologia) Disponível em: https://doi.org/10.1016/j.jacc.2026.01.074
Metadados dos itens 1 a 3 verificados no PubMed; textos integrais dos três consultados no PubMed Central (PMC13019694, PMC13243387, PMC13032450).
Half the Time, Better Performance — and the Effect Size That Is Too Large to Be True
A randomized trial with 124 residents showed that AI-based training cut weekly study from 20.6 to 10.6 hours and improved performance across every domain. That is the headline number. But there is a figure almost no one translates: the equivalent study in obstetrics and gynecology reports an effect size of d = 2.30 — five times the norm in educational research — in a design with no control group. Understanding why an effect that is too large is a warning sign rather than a success is what separates those who consume educational evidence from those who can judge it.
📅 Published on August 11, 2026
Navigate the Article
From the randomized trial to the oversized effect — why judging educational evidence is a clinical competency
- 1. Why This Matters to Anyone Who Teaches Residents
- 2. What the Three Studies Delivered
- 3. The Numbers, Unfiltered
- 4. The Trap: The Effect That Is Too Large
- 5. Instrumentation Is an Educational Act
- 6. The Regulatory Layer — the Gap No One Names
- 7. The Timeline
- 8. What to Do on Monday
- 9. Exercise: What If It Were Cardiology?
- 10. Final Considerations
- 11. References
For non-specialists — the essentials in 60 seconds
This article works on two layers. The technical one, for those who run residency programs or research medical education; and the general one, which requires only curiosity. Every core idea appears twice: once with a hospital example, once with an everyday one.
Medical residency is the in-service training a physician undertakes after graduation, lasting 2 to 6 years, inside a hospital, under supervision. It is where medicine is actually learned — and the scarcest resource there is not information, it is protected time for supervised reasoning.
A randomized controlled trial (RCT) is a design in which participants are allocated by chance to receive the intervention or not. That allocation is what allows us to say the observed difference was caused by the intervention rather than by pre-existing differences between groups.
A pre-post study measures the same group before and after, with no comparison group. It is easier to run and far weaker: anything that happened during that period — including simply maturing — appears blended into the intervention effect.
Effect size (Cohen’s d) is how far apart two groups are, measured in standard deviations. By convention, 0.2 is small, 0.5 medium, 0.8 large. It is the most important word in this article — hold on to it.
Why This Matters to Anyone Who Teaches Residents
Every residency program lives with the same scarcity, and it is not a scarcity of content. Guidelines are public, papers are accessible, video lectures are free. What is missing is protected time for supervised reasoning — the attending sitting beside the resident, looking at the same case, correcting the same error at the moment it happens.
That resource does not scale. One attending can deliver dense, individualized feedback to three, perhaps four residents per rotation. Beyond that, feedback becomes generic, delayed, or simply does not happen. This is why competency-based medical education, adopted as an international standard by the ACGME and by the Royal College of Canada, carries an implementation barrier stated in its own literature: faculty time unavailable for individualized mentoring and assessment.
I taught two AIMED cohorts to obstetrics and gynecology residents in May and June of this year, at the hospital where those residents do their training. They were mixed cohorts, spanning first- to third-year residents in the same room. That experience is what led me to read the trials published in 2026 on precisely this question. What I found was not what I expected — neither in the direction of disappointment nor of confirmation.
What the Three Studies Delivered
Study 1 — the randomized trial (breast oncology, n = 124)
Published on March 12, 2026, in Frontiers in Medicine. One hundred twenty-four breast oncology residents from three urban tertiary teaching hospitals in the same metropolitan region, all nationally accredited centers for standardized residency training. Individual 1:1 randomization by computer-generated sequence prepared by an independent statistician, with allocation concealment via sequentially numbered, opaque, sealed envelopes. Sixty-two in the AIEIT arm, sixty-two in conventional training.
Assessors for the subjective outcomes — clinical reasoning, case presentation, teamwork — were blinded to allocation and were not involved in delivering the intervention. Participants and instructors were not blinded, which is unavoidable in an educational intervention. Prior sample size calculation: expected effect d ≈ 0.6, two-sided alpha 0.05, 80% power, requiring 58 per group. The intervention spanned a three-month full-time clinical rotation.
The intervention had four components: a dynamic knowledge graph using natural language processing over NCCN guidelines, updated monthly; a mixed-reality multidisciplinary team platform with roughly 20 standardized scenarios; a virtual patient–AI mentor system with roughly 25 cases; and a learning analytics dashboard. Declared primary outcome: knowledge acquisition and clinical reasoning performance.
Study 2 — the transesophageal echocardiography simulator (anesthesiology, n = 60)
Published on May 25, 2026, in Frontiers in Surgery. Single-center randomized study at Sichuan Provincial People’s Hospital, between December 2024 and December 2025. Sixty anesthesiology residents, thirty per arm. Both groups received the same instructor-led TEE curriculum — the AI arm additionally trained on a simulator with 3D anatomy visualization, interactive image-reading exercises, and virtual probe manipulation with real-time feedback.
This is the cleanest of the three designs, and it is worth understanding why: because the curricular base is identical, the observed difference is attributable to the addition of the simulator, not to “having technology versus having nothing”.
Study 3 — the obstetrics and gynecology study (n = 28)
Published in 2026 in BMC Medical Education. And here is the difference that changes everything: it is not a randomized trial. It is a prospective interventional study with a pre-post design, with no control group. Twenty-eight residents from three consecutive cohorts — 2021 (n = 10), 2022 (n = 10), 2023 (n = 8) — at a single hospital, the First Affiliated Hospital of Fujian Medical University. The total period was 36 months, and the principal comparison covers 12 months of platform-based training.
Same family of intervention — case simulation platform, adaptive-algorithm personalized pathway, competency-based assessment system. A full rung lower on the evidence hierarchy.
A concrete example: what “half the time” means here
In the oncology trial, the control group studied 20.62 ± 2.48 hours per week. The AI group studied 10.63 ± 2.54 hours — and outperformed them across every domain assessed.
Translated to the residency floor: that is ten hours per week returned to each resident. Over a three-month rotation, roughly 130 hours per person. In a program with twenty residents, 2,600 hours. That is the real asset — not the score gain.
And this is where the number needs Brazilian context: Brazilian medical residency already carries a 60-hour weekly load, of which 10% to 20% must be theoretical activity. Twenty weekly hours of additional study, as in the Chinese control arm, is not the norm in most of our programs — it is above it. The denominator of the gain, therefore, is not the same.
The Numbers, Unfiltered
The table below separates what was measured in a randomized trial from what was measured without a control. That separation is the core content of this article — not the values themselves.
| Outcome | Intervention | Control | Design · Source | Reading |
|---|---|---|---|---|
| Weekly study time (h) | 10.63 ± 2.54 | 20.62 ± 2.48 | RCT · Ji et al. · 2026 | ↑ half the time |
| Molecular classification (% correct) | 93.66 ± 1.92 | 82.65 ± 2.10 | RCT · Ji et al. · 2026 | ↑ p < 0.001 |
| US-guided biopsy — time (min) | 8.5 ± 2.1 | 14.2 ± 3.8 | RCT · Ji et al. · 2026 | ↑ −40% |
| Knowledge retention at 6 months (%) | 86.7 ± 4.2 | 74.2 ± 3.2 | RCT · Ji et al. · 2026 | ↑ sustained |
| Specialty examination pass rate (%) | 90.3 | 67.6 | RCT · Ji et al. · 2026 | ↑ p = 0.0036 |
| Burnout incidence (%) | 35.2 | 61.5 | RCT · Ji et al. · 2026 | ↑ subjective, unblinded |
| Postoperative complications, satisfaction | favorable to the AI arm | RCT · exploratory | ⚠ authors urge caution | |
| TEE — key-view acquisition (%) | 88.2 ± 8.1 | 76.5 ± 9.7 | RCT · Zhang et al. · 2026 | ↑ identical curriculum base |
| TEE — OSATS (0–150) | 121.5 ± 9.2 | 112.4 ± 10.1 | RCT · Zhang et al. · 2026 | ↑ p < 0.001 |
| OB/GYN — clinical reasoning (points) | 68.4 ± 8.1 → 84.9 ± 5.7 · d = 2.36 | pre-post, no control · 2026 | ⚠ see section 4 | |
| OB/GYN — overall competency (points) | 72.9 ± 6.8 → 86.4 ± 4.8 · d = 2.30 | pre-post, no control · 2026 | ⚠ see section 4 | |
The effect size translator — start here
Cohen’s d measures the distance between two means in units of standard deviation. Cohen’s convention: 0.2 small, 0.5 medium, 0.8 large.
In hospital: if you swap one antipyretic for another and mean temperature falls half a degree more, with a standard deviation of one degree, that is d = 0.5. Noticeable, but not transformative.
Outside it: the difference in mean height between adult men and women is approximately d = 2.0. That is a difference you see from across the room, without measuring anyone.
A d = 2.30 in an educational intervention means the average post-training resident outperforms roughly 99% of pre-training residents. In Hattie’s synthesis of more than eight hundred meta-analyses in education, the mean effect of any pedagogical intervention sits around d = 0.40. An effect five to six times larger than the literature average is not grounds for celebration — it is grounds for looking for the alternative explanation.
The Trap: The Effect That Is Too Large
All four domains in the OB/GYN study report Cohen’s d between 1.58 and 2.56. Overall competency: 2.30. Clinical skills: 2.56. Clinical reasoning: 2.36. Professional knowledge: 2.02.
These are enormous effects. And the design that produced them is pre-post with no control group, measuring 12 months of residency.
Ask the obvious question: what else happened to those 28 residents over twelve months? They did their residency. They saw patients, worked shifts, were supervised, read, made mistakes and corrected them. A first-year resident improves substantially from month 1 to month 12 with or without an AI platform. With no control group, there is no way whatsoever to separate the effect of the intervention from the effect of simply having completed another year of residency.
Add to that two well-described mechanisms that systematically inflate pre-post designs: the testing effect — anyone who has taken the assessment once performs better the second time, even without learning anything — and the Hawthorne effect, the performance gain that follows from knowing one is being observed within an innovation project. The authors of the oncology randomized trial explicitly state that they cannot exclude a Hawthorne effect in their own trial, which does have a control group. In the uncontrolled study, the hypothesis cannot even be raised.
None of this means the intervention does not work. It means the OB/GYN study cannot measure how much it works — and that d = 2.30 should be read as the combined effect of intervention, maturation, testing and observation, not as the effect of the platform.
⚠ The caveat that applies equally to the randomized trial
The oncology authors are exemplary in declaring their limits, and they are worth reproducing: the study was conducted only in urban tertiary teaching hospitals, limiting generalizability to resource-limited settings; randomization was not stratified by center, so residual center-level effects cannot be excluded; multiple outcomes were assessed with no formal adjustment for multiple comparisons, increasing type I error risk; some assessments are subjective; and the patient outcomes — postoperative complications, length of stay, satisfaction — are declared exploratory and should be interpreted cautiously.
There is also a structural limitation the authors themselves name and that almost no reader will retain: the intervention is multimodal. Knowledge graph, mixed reality, virtual patient and analytics dashboard were delivered as a package. The observed effect cannot be attributed to AI in isolation — it reflects the combination. The authors correctly call for future factorial designs to isolate each component.
⚠ The transfer barrier to Brazil that no one will notice
All three studies are Chinese. All three. Oncology in Guangzhou, anesthesiology in Sichuan, obstetrics and gynecology in Fujian.
This is not a quality problem — these are well-conducted studies approved by institutional review boards. It is a single-country evidence base problem. Chinese medical residency has its own curricular structure, attending-to-resident ratio, assessment culture and technological infrastructure. None of those variables is neutral with respect to the effect size of an educational intervention.
And there is a second, more uncomfortable layer: in all three cases, the platform evaluated was built by the same group that evaluated it. In the oncology study, the authors write literally “our proposed framework”. This is neither fraud nor bad faith — it is the norm in educational innovation research, where the builders are the testers. But it is precisely the configuration in which the efficacy literature tends to overestimate effects, and it is why independent replication is the mandatory next step, not an optional refinement.
There is not, as far as I was able to verify, a single published Brazilian study with this design. Zero RCTs, zero indexed pre-post studies, in any specialty.
Instrumentation Is an Educational Act
Here the reading turns practical, and the conclusion is counterintuitive.
Look at what is expensive in the randomized trial’s intervention: mixed reality, twenty multidisciplinary team scenarios, twenty-five virtual patients, natural language processing over guidelines refreshed monthly. That is a months-long engineering project with a budget.
Now look at what the authors themselves identify as the mechanism of the effect, in their discussion: personalization of the learning pathway, immediate feedback in a risk-free environment, and the shift from episodic summative assessment to continuous monitoring with targeted remediation. Translated: dense, individualized, high-frequency feedback.
That is exactly what a good attending does. What the platform solves is not the nature of the feedback — it is its scale. And what AI made radically cheaper was not teaching: it was the instrumentation of teaching. Measuring competency before and after, with a validated rubric and paired analysis, costs a spreadsheet and thirty lines of code today.
Apply this to what happens in your own cohorts. Two AIMED cohorts for OB/GYN residents, in May and June, mixed across first to third year — had they been instrumented with a rubric applied on entry and exit, they would have generated the same level of evidence as the Fujian study, with two advantages that study does not have: two independent cohorts rather than three cohorts from the same service, and — more importantly — training year as a clean stratification variable rather than a confounder.
The difference between data and publication, in this case, is not methodological rigor. It is instrumentation — and instrumentation is a decision made before the class starts, not after.
💡 The mixed cohort resolves a confound the published study could not
The most elegant finding in the Fujian study is convergence across cohorts. Before training, the three cohorts had significantly different baseline scores (ANOVA, F = 4.25; p = 0.025) — expected, since they correspond to third, second and first year. Afterwards, the difference vanished (F = 2.14; p = 0.138), with a 47.5% reduction in between-cohort variance (SD 4.04 → 2.12). The authors read this as evidence that the adaptive model meets the distinct needs of distinct stages, standardizing final competency.
The interpretation is plausible. But the design cannot sustain it, and it is worth seeing exactly why: in that study, training year is entirely confounded with calendar period. The 2021 cohort is not merely “third year” — it is a different set of people, recruited at a different moment, with different prior exposure, trained in a different period. There is no way to separate “the training compressed variance across levels” from “these three groups of people simply converged”.
A mixed R1–R3 cohort in the same room eliminates that confound by construction. Same instructor, same day, same content, same case sequence. Training year stops being a confounder and becomes a clean stratification factor — and the question “is the gain larger for those who start further behind?” becomes answerable with a trivial two-way ANOVA.
Add the two cohorts from May and June and you have independent replication of a stratification effect. That is not equivalent to what is published in your specialty. It is methodologically superior — on one specific point, but precisely the one the Fujian authors chose as the most interesting finding of their work.
A concrete example: the stepped wedge, or how to have a control without denying training
The obvious ethical objection to a control group in education is: “I will not deny training to half my cohort.” Correct — and unnecessary.
In a stepped wedge design, every participant receives the intervention, but at different times. Split the cohort into three subgroups and deliver the modules in staggered order: subgroup A gets module 1 in month 1 and module 2 in month 2; subgroup B gets the reverse; and so on.
Each subgroup acts as a control for the others during the period in which it has not yet received that module. No one goes without. And you gain an internal comparator the Fujian study lacks — which would place your cohort, methodologically, above what is published in your specialty.
The Regulatory Layer — the Gap No One Names
CFM Resolution No. 2,454/2026, published in Brazil’s Official Gazette on February 27, 2026, regulates the use of artificial intelligence in medicine. It classifies systems by risk level — low, medium, high or unacceptable — guarantees that the final word on diagnosis, therapy and prognosis always belongs to the physician, prohibits delegating the communication of diagnoses and prognoses to AI, and bars institutions from penalizing physicians who do not follow the machine’s recommendation.
Note what it regulates: clinical decision-making. And note what it does not, by construction, reach: AI applied to medical training. A virtual patient simulator training a resident issues no diagnosis about a real patient. A learning analytics dashboard tracking competency influences no therapeutic conduct. None of the three studies in this article describes a system that CFM 2.454 would classify.
At ANVISA, the Brazilian health regulator, the reasoning arrives at the same place by another route. Regularization of software as a medical device depends on the software aiding diagnosis, monitoring a clinical condition, or influencing therapeutic decisions. An educational platform operating on simulated cases does not meet that trigger.
The correct reading of this is not “it is unregulated, do as you please”. It is the opposite: this is simultaneously the fastest-adopting and least-regulated area of medical AI. When a residency program adopts a platform that decides what each resident studies and how each resident is assessed, no CFM resolution and no ANVISA registration demands prior validation. The only barrier is the quality of evidence the program itself requires before adopting — and that barrier is institutional, voluntary, and today all but nonexistent.
⚠ Where the boundary becomes regulated again
There is a point at which the argument above stops holding, and it is worth marking: the moment an educational platform begins operating on real patient data — medical records, imaging from the service itself, clinical outcomes attributed to a specific resident — it leaves the purely educational domain. At that point, depending on the case, Brazil’s data protection law, mandatory research ethics committee approval, and, depending on function, the software-as-a-medical-device discussion all come into play.
The oncology trial handled this correctly: the adaptive algorithms were trained exclusively on aggregated, anonymized learner–content interaction data, with no access to patient data or electronic health records, and the system generated no autonomous diagnostic or therapeutic recommendations. That boundary is a design decision, not an implementation detail.
The Timeline
2021–2024 · The pre-evidence phase
AI in medical education appears as isolated, fragmented tools. Virtual reality simulation, virtual patients and automated assessment exist, but without comparative design. The literature is about feasibility and acceptance, not efficacy.
2026 · The first body of randomized evidence — and where you are now
Three studies, three specialties, one country. Two randomized, one without a control. No independent replication, no data outside China, no Brazilian study. The effects are large; so is the uncertainty about how much of them is real.
Next 12–24 months · The replication window
This is the interval in which factorial designs will begin isolating which component carries the effect, and in which replications outside China will test whether the magnitude holds. It is also the window in which a well-instrumented Brazilian cohort enters the literature as a first, not as belated confirmation.
Beyond · Curricular consolidation
Once the evidence stabilizes, the discussion migrates from “does it work?” to “how much of the theoretical workload can be redesigned?”. At that point the decision stops being pedagogical and becomes training policy — national residency commissions, program committees, faculty boards.
What to Do on Monday
If you run or teach in a residency program
- ☐ Before adopting any platform, ask about the design of the study behind it. If the answer is “we improved competency by X%”, ask: compared to whom? The absence of a control group is the most important piece of information in the sales material — and it is always the piece that is not in it.
- ☐ Be suspicious of effect sizes above d = 1.0 in educational interventions. The literature norm is d ≈ 0.4. An effect far above that demands an alternative explanation before celebration.
- ☐ Ask who built the platform being evaluated. If the publishers are the developers, the finding remains valid — but it requires independent replication before becoming the basis for an institutional decision.
- ☐ Instrument your next cohort before it starts. A four-domain competency rubric, applied on entry and exit, on the same scale. Cost: a spreadsheet. Without it, the cohort’s data is lost and does not come back.
- ☐ If you want a control without denying training, use a stepped wedge. Staggered module order across subgroups delivers an internal comparator at zero ethical cost.
- ☐ If your cohort is mixed by training year, record each participant’s year. It is the cheapest variable to collect and the highest-yield analytically: with delivery held constant, it lets you test whether the gain is larger for those who start further behind — a question the published literature cannot answer without confounding.
- ☐ Record study time, not only performance. The most robust finding in the randomized trial is the reduction in workload with performance preserved or improved. It is the outcome with the highest institutional value and the easiest to measure.
- ☐ Submit to the ethics committee before collecting. Educational data from identifiable participants requires prior approval for publication. It is the step that most often retrospectively kills an otherwise good cohort.
Exercise: What If It Were Cardiology?
The pattern in the three studies transfers, but not uniformly. What changes from specialty to specialty is not the technology — it is which competency is the bottleneck and which is measurable with an objective rubric. The exercise is worth doing out loud, because it exposes the design reasoning better than any ready-made conclusion.
Cardiology: the most favorable case of all, and almost no one has noticed
Cardiology has a property that oncology and gynecology do not: three core competencies with an objective, digitizable and scalable gold standard.
Electrocardiography. It is the only high-volume clinical competency in which the assessment item is a file, the answer key is consensual, and the case bank can hold tens of thousands of examples with known outcomes. Nothing equivalent exists in gynecology. An adaptive system serving the next tracing calibrated by the resident’s previous error is, technically, the simplest application in this entire article — and the one with the highest repetition density.
Echocardiography. Here direct, favorable evidence already exists, albeit outside cardiology: the transesophageal echocardiography trial with sixty anesthesiology residents showed gains in key-view acquisition (88.2% versus 76.5%) and reduced procedure time (27.1 versus 32.4 minutes), with an identical curricular base in both arms. Transthoracic echocardiography in cardiology residency has the same profile: view acquisition is objectively measurable, and OSATS is a validated, transferable instrument.
Hemodynamics and electrophysiology. These are the domains in which real exposure is scarcest and most dependent on institutional volume — precisely the condition in which a simulator delivers the greatest marginal value. And it is where the competency reference has just been updated: the 2026 ACC/AHA/HRS Advanced Training Statement on Clinical Cardiac Electrophysiology revised training requirements and designated competencies with explicit alignment to cutting-edge technologies.
The detail that inverts the order of priority
Intuition says “start with the ECG, it is the easiest”. Reading the three studies suggests the opposite.
Note which domain gained most in the OB/GYN study: clinical reasoning, which started from the lowest baseline score (68.4) and rose the most (16.5 points). The authors record that higher-order cognitive skills — risk assessment accuracy and comprehensive differential diagnosis formulation — had the lowest initial scores and the greatest improvement margins, 27.4% and 24.8%. The same pattern appears in the oncology trial, whose most pronounced gains were in specialized oncology reasoning and multidisciplinary collaboration, not in isolated technical tasks.
In other words: the effect is largest where the competency is most complex and least trainable by repetition, not where it is easiest to digitize. In cardiology, that shifts the bet from the ECG in isolation toward decision-making under uncertainty — the syncope case discussion, the device indication, preoperative risk stratification, the management of intermediate-risk chest pain. That is the territory where variability between cardiologists is greatest and the rubric is hardest to write. It is also where the gain, on the available data, would be largest.
This is an uncomfortable result for curriculum planners, because it inverts the order of least effort. But it is what the two studies with domain-level data show.
The question to take to your faculty board
Regardless of specialty, the exercise reduces to three questions, in this order:
1. Which competency in my program has the lowest baseline score and the greatest variability between residents? (That is where the gain will be largest.)
2. Does that competency have a consensual answer key among my attendings? (If not, the problem is not the platform — it is the rubric, and it must be solved first.)
3. Can I measure it with the same ruler on entry and exit? (If not, there is no study — there is an impression.)
If the answer to question 2 is no, you have just discovered something more valuable than any platform: your team does not agree on what correct performance is. No AI system solves that, and every AI system will silently amplify it.
Learn to Judge the Evidence Before Adopting the Tool
The AIMED methodology trains physicians who build — not merely consume — clinical AI tools. Critical appraisal of study design, effect size under weak designs, instrumentation of your own cohort, and the regulatory boundary between education and clinical decision-making are part of the curriculum, because they are the same competency seen from different angles.
Final Considerations
Ji and colleagues’ trial is good. Computer-generated randomization, allocation concealment by opaque envelope, assessors blinded to subjective outcomes, prior sample size calculation and a declared primary outcome. Halving weekly study time while still improving performance is a finding that deserves attention, and the authors report their own limitations with above-average honesty.
What deserves criticism is not the studies — it is the reading habit that will form around them. “AI improves resident training” will circulate as established fact, resting on three papers from a single country, two of them with the same group evaluating its own platform, one of them without a control group and reporting an effect five times larger than the educational literature average.
💡 Connecting the Dots: the authority asset here is not the “half the time” — it is the fact that the effect size of an educational intervention is inversely proportional to the quality of the design that measured it, and the entire field still reads a large effect as good news. The randomized trial, with a control, reports differences that are substantial but plausible; the uncontrolled study reports d = 2.30, and it is precisely the weakest design that produces the most impressive number. That is no coincidence: it is the predictable signature of a design that adds intervention, maturation, testing and Hawthorne effect into a single number and calls all of it the platform’s effect. But the second leap is the one almost nobody makes. If the effect is largest where the competency is most complex and least trainable by repetition — and the two studies with domain-level data show exactly that — then the right target in any specialty is not the task that is easiest to digitize; it is decision-making under uncertainty, which is where the rubric is hardest to write because the attendings themselves do not agree. In other words: the barrier to instrumenting medical training was never technological. It is that instrumentation forces a service to declare, in writing and before the class begins, what it considers correct performance — and that is a conversation most programs have never had. Whoever has that conversation first does not gain a better platform: they gain the ruler. And whoever defines the ruler defines what the next generation of the specialty will consider competence.
References
- Ji F, Xiao W, Li X. AI-driven intelligent training enhances clinical competence in oncology residency: a randomized controlled trial. Front Med (Lausanne). 2026 Mar 12;13:1768388. (RCT, n = 124; weekly study time 10.63 ± 2.54 h vs 20.62 ± 2.48 h; 6-month retention 86.7 ± 4.2% vs 74.2 ± 3.2%; exam pass rate 90.3% vs 67.6%, p = 0.0036; burnout 35.2% vs 61.5%) Available at: https://doi.org/10.3389/fmed.2026.1768388
- Zhang M, Song SY, Zhou H, Xu G, Tong X, Fan D, Lei Q. Application of an artificial-intelligence-based transesophageal echocardiography simulation system in residency training. Front Surg. 2026 May 25;13:1762740. (single-center RCT, n = 60; key-view acquisition 88.2 ± 8.1% vs 76.5 ± 9.7%; OSATS-TEE 121.5 ± 9.2 vs 112.4 ± 10.1; time 27.1 ± 3.9 vs 32.4 ± 4.6 min) Available at: https://doi.org/10.3389/fsurg.2026.1762740
- Construction and implementation of an AI-enhanced progressive training model for clinical competency development in obstetrics and gynecology residents. BMC Med Educ. 2026. (pre-post study without control group, n = 28, three cohorts, single center; overall competency 72.9 ± 6.8 → 86.4 ± 4.8, d = 2.30; clinical reasoning d = 2.36) Available at: https://doi.org/10.1186/s12909-026-08751-5
- Conselho Federal de Medicina. Resolution CFM No. 2,454 of February 11, 2026 — Regulating the use of artificial intelligence in medicine. Official Gazette of Brazil, February 27, 2026; Ed. 39, Section 1. Available at: https://sistemas.cfm.org.br/normas/arquivos/resolucoes/BR/2026/2454_2026.pdf
- American College of Cardiology / American Heart Association / Heart Rhythm Society. 2026 ACC/AHA/HRS Advanced Training Statement on Clinical Cardiac Electrophysiology. J Am Coll Cardiol. 2026. Available at: https://doi.org/10.1016/j.jacc.2026.01.074
Metadata for items 1 to 3 verified on PubMed; full texts of all three consulted on PubMed Central (PMC13019694, PMC13243387, PMC13032450).

