Brief
AWS publishes 38 open-source agent skills aimed at HCLS reasoning errors
An AWS machine learning blog post introduces 38 open-source agent skills across 11 healthcare and life sciences domains, aimed at agents that cite the correct guideline but apply it incorrectly. The evidence given is a headline summary: installation steps, three worked use cases and a 410-prompt evaluation with a 70-86% win rate.
An AWS machine learning blog post describes 38 open-source agent skills spanning 11 healthcare and life sciences domains, aimed at a specific failure: agents on foundation models citing the right clinical or scientific guideline and then misapplying its decision framework. The post provides installation steps, three worked use cases and a 410-prompt evaluation reporting a 70-86% win rate.
The evidence stops there. The summary does not name the models tested, the baseline the skills were measured against, or how prompts were scored, so the win rate cannot yet be read as a general claim about agent reasoning. For teams already running agents over clinical or scientific material, the skills are the concrete artefact; for everyone else the open question is whether the gain survives outside that evaluation set.
Our reading
Most agent failures in specialised work are not retrieval failures but application failures, and packaging decision frameworks as installable skills is a repair pattern worth watching beyond healthcare and life sciences. Implementation and evaluation teams running agents over regulated or technical material should care, because the artefact is inspectable and the evaluation is described rather th…
What to do or watch
Read the installation steps and run one skill against prompts your own agents already fail on, scoring outputs the same way you would score a human reviewer; the unresolved question is which models, baselines and scoring rules produced the reported 70-86% win rate.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by aws.amazon.com
- The post shares 38 open-source agent skills across 11 HCLS domains.
- AI agents on foundation models often misapply healthcare and life sciences decision frameworks, citing the right guideline but applying it incorrectly.
- The post includes installation steps and three worked use cases.
- A 410-prompt evaluation showed a 70-86% win rate.
Sources
- AWS Machine Learning BlogText stored 16 September 2026
How this story was checked. Written from the 1 page listed above, stored 16 September 2026; claims checked against that stored text on 16 September 2026.
What that means
- 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.