Extracting Clinical Entities
SkillFiles & storageLets your agent find medical terms like diseases, drugs, and genes in text and return them as JSON, CSV, or HTML.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Extracting Clinical Entities skill
About this capability
Run clinical and biomedical named-entity recognition on medical text with OpenMed's analyze_text. Use when the user wants to extract diseases, drugs, anatomy, genes, or other biomedical entities from notes; needs NER output as dict/json/html/csv; wants to filter by confidence, group entities, toggle
What this skill tells your AI
The instructions your AI receives, as published by maziyarpanahi/openmed in skills/extracting-clinical-entities/SKILL.md and read by ahel’s review.
openmed.analyze_text runs a token-classification model over medical text and
returns structured entities with character offsets and confidence scores. It runs
on-device after a one-time model download.
When to use
- Pull diseases, medications, anatomy, genes, proteins, etc. out of clinical text.
- You need exact character spans (start/end) plus confidence per entity.
- You want output as objects, JSON, an HTML highlight view, or CSV.
- You are building the "extract entities" stage of a clinical NLP pipeline.
To choose a model, see choosing-openmed-models. To load it once and reuse it,
see loading-openmed-models. In a PHI workflow, de-identify first (see
deidentifying-clinical-text), then run NER on the redacted text.
Install
pip install "openmed[hf]"
Quick start
import openmed
note = (
"Patient prescribed 500 mg metformin for type 2 diabetes mellitus. "
"Reports intermittent chest pain; ruled out myocardial infarction."
)
result = openmed.analyze_text(
note,
model_name="disease_detection_superclinical", # registry key, HF id, or local path
output_format="dict", # dict | json | html | csv
confidence_threshold=0.5,
)
for ent in result.entities:
print(f"{ent.label:12} {ent.text!r:40} {ent.confidence:.2f} [{ent.start}:{ent.end}]")
With output_format="dict" you get a PredictionResult. The fields you use most:
result.text # the original input text
result.entities # list of entity objects
result.model_name # which model produced these
ent.text # the surface string
ent.label # entity type, e.g. "DISEASE"
ent.confidence # model score in [0, 1] (NOTE: .confidence, not .score)
ent.start / ent.end # character offsets into result.text
Output formats
analyze_text(...) returns different types depending on output_format:
output_format | Return type | Use for |
|---|---|---|
"dict" (default) | PredictionResult object | Programmatic access via .entities. |
"json" | str (JSON) | Logging, APIs, writing to disk. |
"html" | str (HTML) | A highlighted preview of the note. |
"csv" | str (CSV) | Spreadsheet / quick review. |
import openmed
note = "Started atorvastatin 40 mg; history of myocardial infarction."
json_str = openmed.analyze_text(note, output_format="json")
html_str = openmed.analyze_text(note, output_format="html") # render in a browser
csv_str = openmed.analyze_text(note, output_format="csv")
Key parameters
openmed.analyze_text(
text,
model_name="disease_detection_superclinical",
output_format="dict",
confidence_threshold=0.5, # drop entities below this score; None keeps all
aggregation_strategy="simple", # HF subword aggregation; None for raw tokens
group_entities=False, # merge adjacent same-label spans into one
include_confidence=True, # include scores in formatted output
sentence_detection=True, # pySBD sentence splitting (better long-doc spans)
sentence_language="en",
loader=None, # pass a reused ModelLoader (see loading skill)
)
confidence_threshold— the most useful knob. Use the model'srecommended_confidence(fromget_model_info) as a starting point.group_entities=True— merges"type","2","diabetes"fragments into a single"type 2 diabetes"span. Turn on for cleaner output.sentence_detection=True(default) — splits long notes into sentences before inference for more accurate offsets and to respect model max length. Requires pySBD; if unavailable it silently falls back to whole-text inference.
Save results to JSONL
One line per note keeps offsets and labels for downstream grounding or eval:
import json
import openmed
notes = [
"Type 2 diabetes managed with metformin.",
"Acute myocardial infarction; started aspirin and atorvastatin.",
]
with open("entities.jsonl", "w", encoding="utf-8") as fh:
for i, note in enumerate(notes):
result = openmed.analyze_text(note, output_format="dict")
fh.write(json.dumps({
"doc_id": i,
"text": result.text,
"model": result.model_name,
"entities": [
{"label": e.label, "text": e.text,
"start": e.start, "end": e.end,
"confidence": round(e.confidence, 4)}
for e in result.entities
],
}) + "\n")
Store offsets and labels, not extra copies of free text, in PHI contexts.
CLI
openmed analyze --text "Type 2 diabetes managed with metformin." \
--model disease_detection_superclinical \
--format json \
--threshold 0.5 \
--group
# Or analyze a file:
openmed analyze --input-file note.txt --model disease_detection_superclinical -o csv
Flags: --text/-t, --input-file/-f, --model/-m, --format/-o
(dict|json|html|csv), --threshold/-c, --group, --no-confidence,
--sentence-detection/--no-sentence-detection.
Hand-off to / from OpenMed
-
From
loading-openmed-models: pass your reusedloader=so a batch loads weights once. -
From
deidentifying-clinical-text: run NER onresult.deidentified_text, not raw PHI:deid = openmed.deidentify(raw_note, method="mask", policy="hipaa_safe_harbor") ner = openmed.analyze_text(deid.deidentified_text, output_format="dict") -
To terminology grounding (out-of-process): map
ent.text/ent.labelto RxNorm / LOINC / SNOMED using the user's own licensed service — OpenMed does not bundle restricted terminologies. -
To batch processing: for large corpora use
openmed.process_batch(...)/BatchProcessor(see processing utilities) with a shared loader.
Edge cases & gotchas
- Attribute is
.confidence, not.score. Entity objects extendEntityPrediction(text,label,confidence,start,end). - Right model for the labels. A Disease model won't emit oncology staging or
gene labels — pick the category in
choosing-openmed-modelsand checkentity_types. - Offsets index
result.text. Slice the original string withstart:end; the surface form inent.textis whitespace-trimmed. - Long documents: keep
sentence_detection=Trueso chunks respect the model's max length (get_model_max_length); disabling it can truncate long notes. - NER assists, it does not diagnose. Treat output as decision support; surface a disclaimer for any clinical-facing use.
- No raw PHI in logs. Log labels, offsets, and hashes — never patient text.
Standards & references
- Token classification / NER (Hugging Face): https://huggingface.co/docs/transformers/tasks/token_classification
- pySBD sentence boundary detection: https://github.com/nipunsadvilkar/pySBD
- OpenMed model cards: https://huggingface.co/OpenMed
Signals
- GitHub stars
- 5k
- Forks
- 666
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
extracting-clinical-entities- Source
- github.com/maziyarpanahi/openmed