Instructions to use ChengLi0228/Angel-Observer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ChengLi0228/Angel-Observer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ChengLi0228/Angel-Observer") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ChengLi0228/Angel-Observer") model = AutoModelForCausalLM.from_pretrained("ChengLi0228/Angel-Observer", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ChengLi0228/Angel-Observer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ChengLi0228/Angel-Observer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChengLi0228/Angel-Observer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ChengLi0228/Angel-Observer
- SGLang
How to use ChengLi0228/Angel-Observer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ChengLi0228/Angel-Observer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChengLi0228/Angel-Observer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ChengLi0228/Angel-Observer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChengLi0228/Angel-Observer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ChengLi0228/Angel-Observer with Docker Model Runner:
docker model run hf.co/ChengLi0228/Angel-Observer
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("ChengLi0228/Angel-Observer")
model = AutoModelForCausalLM.from_pretrained("ChengLi0228/Angel-Observer", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))Angel-Observer
Angel-Observer reads a patient's presenting complaints and builds a directed symptom network (a Network Model): the thoughts, emotions, behaviors, bodily sensations and events in the case, and which ones set off which. It is stage 1 of Angel, a two-stage simulator of psychotherapy patients; Angel-Actor then role-plays the patient.
🎮 Live demo · 💻 Code: ANGEL-UserSim · 🎠Angel-Actor
short description ──► Angel-Observer ──► long profile + symptom network ──► Angel-Actor ──► patient replies
What it does
One checkpoint answers two prompts:
| Step | Input | Output |
|---|---|---|
| S1: nodes | presenting complaints | <think>…</think><GRAPH>{"symptoms": [...], "external_factors": [...]}</GRAPH> |
| S2: edges | complaints + the S1 nodes | <think>…</think><GRAPH>{"links": [{"from": ..., "to": ...}]}</GRAPH> |
In the Angel pipeline it also expands a short patient description into a detailed long profile for the Actor.
How to use
The system prompts the model was trained on are in the code repository, so the easiest way is through it:
git clone https://github.com/Scarelette/ANGEL-UserSim.git && cd ANGEL-UserSim
pip install -r model_training/observer/requirements-observer.txt
# one symptom network per case (JSONL with a "Complaints" field)
python -m model_training.observer.predict_network \
--input data/examples/observer/case_reports.jsonl --output networks.jsonl
The model is downloaded from this page on first use. From Python:
from model_training.observer.predict_network import load_observer, predict_network
tokenizer, model = load_observer("ChengLi0228/Angel-Observer")
net = predict_network(tokenizer, model, "Alex, 29, reports three months of low mood, "
"poor sleep and withdrawing from friends after being laid off ...")
print(net["symptoms"])
print(net["graph"]) # [{"from": "being laid off", "to": "low mood"}, ...]
For the full Angel patient (Observer + Actor) see
model_usage/.
It uses the Qwen3 chat template, sampling with temperature 0.1 and top-p 0.9,
and a bf16 model needs about 17 GB of GPU memory.
Training
Qwen3-8B, fine-tuned in four stages (QLoRA adapters on a 4-bit copy, each merged back in bf16):
| Stage | Data | Settings |
|---|---|---|
| SFT, S1 (nodes) | 5,599 rows | LoRA r=64 / α=16, lr 1e-4, 3 epochs |
| GRPO, S1 | 5,089 prompts, 600 steps | reward 0.4 · format + 0.6 · node recall (MiniLM cosine ≥ 0.75 against reference nodes) |
| SFT, S2 (edges) | 2,138 rows | LoRA r=64 / α=16, lr 1e-4, 3 epochs |
| GRPO, S2 | 800 steps | reward 0.1 · format + 0.6 · edge precision + 0.3 · node coverage − size penalty; edges scored by a gpt-5-mini judge |
The training data were built with GPT-5 from 510 of the 516 published
psychotherapy case reports in PSYCHE, our graph-grounded dataset for
psychological user simulation (release coming soon). The case reports are
copyrighted and are not redistributed. Full recipe:
model_training/observer/.
Intended use and limitations
- Research use in user simulation, for example generating varied, structured patient cases to train or evaluate therapy-support systems and clinicians-in-training.
- Not a diagnostic or clinical tool. The networks are the model's reading of a text, not clinical assessments, and they can be wrong, incomplete or reflect biases in the source case reports.
- Trained on English case-report language; other languages and very different writing styles are untested.
- Case material can involve self-harm, suicidality, trauma and substance use, so the model's outputs can too.
Citation
Coming soon.
- Downloads last month
- 45
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ChengLi0228/Angel-Observer") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)