--- license: apache-2.0 base_model: - Qwen/Qwen3.5-4B library_name: transformers pipeline_tag: text-generation tags: - medical - agent - tool-use - web-search - reinforcement-learning --- # MedSearch-R1 MedSearch-R1 is a locally deployable medical search agent policy initialized from Qwen3.5-4B. The model is trained through cold-start knowledge distillation, step-level on-policy distillation, trajectory-level on-policy distillation, and accuracy-based reinforcement learning. This repository contains the policy-model weights and tokenizer only. The Search--Visit tools, source-policy filters, helper-model configuration, and evaluation pipeline are not embedded in the checkpoint. Exact agent-loop code and reproducibility configurations will be provided in the associated GitHub repository. ## Loading ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "xiaohan666/MedSearch-R1" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto", ) ``` The checkpoint uses the Qwen3.5 architecture and requires a Transformers release with `Qwen3_5ForCausalLM` support. ## Intended Use The model is intended for research on medical reasoning agents and tool-augmented language models. Reproducing the paper's agent results requires the accompanying Search--Visit loop and source-filtering configuration. ## Limitations MedSearch-R1 is not a medical device and must not be used as a substitute for professional medical judgment. Generated answers and retrieved evidence can be incomplete or incorrect. Local policy inference reduces full-context exposure to external model providers, but generated search queries may still reveal medical concepts and do not constitute a formal privacy guarantee. ## License The model is released under the Apache License 2.0, following the license of the Qwen3.5-4B base model.