Official code for CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework, published at ICLR 2026.
CARE answers medical visual questions through three specialist stages: question-conditioned entity proposal, entity-referring segmentation, and evidence-grounded VQA (EG-VQA). CARE-Flow runs global, mask, and zoom-in evidence and votes over the three answers. CARE-Coord lets a coordinator choose evidence and review the expert model's reasoning before returning an answer.
CARE is research software. It is not a medical device and must not be used for diagnosis, treatment, or clinical decision-making.
| Area | Implementation |
|---|---|
| Entity proposal | InternVL3-2B DAPO with semantic entity matching, count, repetition, and format rewards |
| Referring segmentation | SA-Med2D with BioClinicalBERT text conditioning and entropy-based mask confidence |
| EG-VQA | InternVL3-2B and 8B SFT followed by DAPO over global, binary-mask, and zoom-in clues |
| Inference | OpenAI-compatible endpoints for the entity and EG-VQA models, local SA-Med2D, majority vote, GPT-5 coordination, and a separately labeled local selector |
| Evaluation | Closed-answer accuracy, GPT-4o open-answer judging, and baseline adapters |
This release contains source code and prepared JSONL annotations. Pretrained CARE weights are not available at this time. Dataset images are obtained separately from their original providers. See docs/ARTIFACTS.md for the expected local artifact layout and docs/INSTALL.md for dependencies. The training commands below can be used to produce local checkpoints.
Run from the repository root:
conda create -y -n care-core -c conda-forge --override-channels python=3.11 pip
conda activate care-core
python -m pip install -e ".[test]"
pytest -q -sTraining and inference use different framework versions. Create the task-specific environments in docs/INSTALL.md instead of installing all requirement files into one environment.
The repository includes the prepared VQA train/test annotations (including the exact OmniMedVQA 4k/3k subsets) and historical entity-training annotations. annotations/README.md lists the original image sources, exact image-directory mappings, record counts, and checksums.
Download and link the original images as described there, then validate every image reference and install the annotations at the paths used by the configs:
python scripts/prepare_annotations.py
# Or select only the datasets you have downloaded:
python scripts/prepare_annotations.py --dataset vqa_rad slakeThe release includes an explicit mask-annotation stage. It calls the entity endpoint, runs the local segmenter, records each mask confidence, then builds global, mask, and zoom rows for every source row with a non-empty answer. The canonical SLAKE config reports and skips the one blank-answer row in the audited archive.
Generated evidence is isolated by output dataset under
data/processed/evidence/<output-stem>/<record-id>/; identical record IDs in
different datasets therefore cannot overwrite one another. Advanced configs may
set evidence_root and evidence_namespace explicitly.
care data prepare --config configs/data/annotate_vqa_rad.yaml
care data prepare --config configs/data/annotate_slake.yaml
care data prepare --config configs/data/annotate_omnimedvqa.yaml
care data prepare --config configs/data/build_evidence_vqa_rad.yaml
care data prepare --config configs/data/build_evidence_slake.yaml
care data prepare --config configs/data/build_evidence_omnimedvqa.yamlNormalize a final evaluation split from its legacy InternVL row format before inference:
care data prepare --config configs/data/normalize_vqa_rad_test.yaml
care data prepare --config configs/data/normalize_slake_test.yaml
care data prepare --config configs/data/normalize_omnimedvqa_test.yaml
care data prepare --config configs/data/normalize_vqamed2019_test.yamlThese commands preserve the reference answer and evaluation method, and fail if any referenced image is missing. OmniMedVQA's Images/... paths are rebased to OmniMedVQA/Images/... under data/.
The historical entity-training JSONL can be installed after linking its SA-Med2D-20M images. Read its quality notes before training; the snapshot is preserved without filtering.
python scripts/prepare_annotations.py --dataset entityAlternatively, synthesize fresh entity-proposal supervision from SA-Med-20M metadata:
care data prepare --config configs/data/entity_synthesis.yamlThe exact schemas, record counts, and data-quality notes are in docs/DATA.md.
Every command supports --dry-run. Dotted overrides use --set KEY=VALUE.
care train --config configs/train/entity_dapo.yaml
care train --config configs/train/segmenter.yaml
care train --config configs/train/egvqa_sft_2b.yaml
care train --config configs/train/egvqa_dapo_2b.yamlReplace 2b with 8b for the larger EG-VQA model. The canonical configs enforce a global batch of 64 and the paper's DAPO settings. docs/REPRODUCIBILITY.md maps every setting to the paper and retained development code.
DAPO produces a PEFT adapter. Merge it into the base or SFT checkpoint before standard vLLM serving:
python scripts/merge_lora.py \
--base-model OpenGVLab/InternVL3-2B-Instruct \
--adapter outputs/entity-dapo \
--output outputs/entity-dapo-merged
python scripts/merge_lora.py \
--base-model outputs/egvqa-sft-8b \
--adapter outputs/egvqa-dapo-8b \
--output outputs/egvqa-dapo-8b-mergedSet a non-secret local API key consistently in vLLM and the CARE configs. The model names exposed by the servers must match the model values in configs/inference/*.yaml.
export CARE_API_KEY=local-development-key
vllm serve outputs/entity-dapo-merged \
--served-model-name care-entity-proposal --port 8001 --api-key "$CARE_API_KEY" \
--trust-remote-code --limit-mm-per-prompt 'image=8'
vllm serve outputs/egvqa-dapo-8b-merged \
--served-model-name care-egvqa --port 8002 --api-key "$CARE_API_KEY" \
--trust-remote-code --limit-mm-per-prompt 'image=8'The serving commands above use locally trained and merged checkpoints. Set
models.segmenter.checkpoint in the inference config to your trained SA-Med2D checkpoint
(or place it at the default checkpoints/care-segmenter.pth). Keep the served model
names care-entity-proposal and care-egvqa consistent with the inference configs.
Run the fixed workflow or a coordinator in another shell:
care infer --config configs/inference/care_flow.yaml --set input=data/evaluation/vqa_rad.jsonl
care infer --config configs/inference/care_coord.yaml --set input=data/evaluation/vqa_rad.jsonl
care infer --config configs/inference/care_coord_local.yaml --set input=data/evaluation/vqa_rad.jsonl
care evaluate --config configs/evaluation/closed.yaml
care evaluate --config configs/evaluation/open.yamlInference output retains the input question, answer, image path, and metadata alongside the CARE trace. When a mask or zoom request degrades to the global image, the effective fallback reason is recorded in the trace's fallbacks field. This makes the evaluation configs usable without a separate join step.
pytest -q -s
python scripts/check_release.py
bash scripts/dry_run_pipeline.shGPU tests are opt-in because they require checkpoints:
CARE_RUN_GPU_TESTS=1 pytest -q -s -m gpuCARE's original code is Apache-2.0. The source release also contains vendored
MIT, Apache-2.0, and BSD-2-Clause components. Their complete texts and component
mapping are in third_party/licenses/ and
THIRD_PARTY_NOTICES.md. Dataset and checkpoint
licenses are separate and must be verified before those artifacts are
redistributed. See PROVENANCE.md and
CITATION.cff.
@inproceedings{du2026care,
title={CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework},
author={Du, Yuexi and Wang, Jinglu and Liu, Shujie and Dvornek, Nicha C. and Lu, Yan},
booktitle={International Conference on Learning Representations},
year={2026},
eprint={2603.01607},
archivePrefix={arXiv}
}