license: apache-2.0
base_model: Qwen/Qwen2.5-7B-Instruct
library_name: peft
tags:
- lora
- peft
- qwen2.5
- multi-agent
- cooperative
- overcooked
language:
- en
DriCo: Planner / Actor / Coordinator LoRA Adapters
Four LoRA adapters fine-tuned on Qwen/Qwen2.5-7B-Instruct for a multi-agent cooperative-cooking benchmark.
A single base model is loaded once at inference; the three roles are switched by activating the matching adapter.
| Folder | Role | Training |
|---|---|---|
planner/ |
High-level subgoal generation in natural language (e.g. "Pick up carrot and put it in pot0.") | SFT |
actor/ |
Low-level action chunking β turns a subgoal into up to 3 ml_actions |
SFT |
coordinator/ |
Shared-context generation + PASS/REJECT critique of planner subgoals | SFT |
coordinator_dpo/ |
Coordinator further refined with DPO | SFT β DPO |
Quick start
End-to-end: clone the project repo, drop these adapters into the right folder, run eval.
# 1) Clone the project repo
git clone https://github.com/anonymous-projectpage/DriCo.git
cd DriCo
# 2) Install deps (in a fresh env)
pip install -U huggingface_hub
pip install -r requirements.txt # if the repo has one
# 3) Download these adapters into drico/out/
hf download anonymous-24421/DriCo-adapters \
--local-dir drico/out
# 4) (Optional) download the training/DPO data into drico/dataset/
hf download anonymous-24421/DriCo \
--repo-type dataset \
--local-dir drico/dataset
# 5) Run evaluation
cd drico
CUDA_VISIBLE_DEVICES=0 python eval_proposed.py \
--base_model Qwen/Qwen2.5-7B-Instruct \
--adapter_root out/ \
--layout circuit_env \
--recipe_split test \
--statistics_save_dir data/eval_test
After step 3, the directory tree will look like:
DriCo/
βββ drico/
βββ out/
βββ README.md β this file
βββ planner/
β βββ adapter_config.json
β βββ adapter_model.safetensors
β βββ ...
βββ actor/
β βββ ...
βββ coordinator/
β βββ ...
βββ coordinator_dpo/
βββ ...
That's exactly what eval_proposed.py --adapter_root out/ expects.
Selective download
Only the actor adapter:
hf download anonymous-24421/DriCo-adapters \
--include "actor/*" \
--local-dir drico/out
Only the DPO-refined coordinator:
hf download anonymous-24421/DriCo-adapters \
--include "coordinator_dpo/*" \
--local-dir drico/out
Manual loading in Python
If you want to use the adapters outside the eval script:
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "Qwen/Qwen2.5-7B-Instruct"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(
base, torch_dtype="bfloat16", device_map="auto",
)
# Register the three roles as named adapters on one base model
model = PeftModel.from_pretrained(
model, "anonymous-24421/DriCo-adapters",
subfolder="planner", adapter_name="planner",
)
model.load_adapter("anonymous-24421/DriCo-adapters",
subfolder="actor", adapter_name="actor")
model.load_adapter("anonymous-24421/DriCo-adapters",
subfolder="coordinator_dpo", adapter_name="coordinator")
# Switch role per call
model.set_adapter("planner")
out = model.generate(
**tok("Plan the next subgoal: ...", return_tensors="pt").to(model.device),
max_new_tokens=50,
)
print(tok.decode(out[0], skip_special_tokens=True))
model.set_adapter("actor")
# ... generate the next action chunk ...
Swap subfolder="coordinator_dpo" for subfolder="coordinator" to use the
SFT-only coordinator instead of the DPO-refined one.
Related
- Project repo: https://github.com/anonymous-projectpage/DriCo
- Training data: https://huggingface.co/datasets/anonymous-24421/DriCo
- Base model: https://huggingface.co/Qwen/Qwen2.5-7B-Instruct
Intended use
Research on multi-agent cooperative task planning, role-conditioned generation, and critiqueβrefine pipelines. Not intended for general chat/assistant use.
Limitations
- Trained for a specific cooperative-cooking environment with a fixed action
vocabulary (
pickup,put_obj_in_utensil,cook,bake,stir, ...). - Outputs outside that domain are not guaranteed to be coherent.
- The coordinator's critique is a heuristic signal, not a guarantee of task correctness.
Citation
@misc{drico2026,
title = {DriCo: Planner / Actor / Coordinator for Cooperative Cooking},
author = {Anonymous},
year = {2026},
note = {Under review}
}