Curious about the longbenchv2 benchmark

#15
by Vincent-luo - opened

Thanks for your great work!
I am interested in the long context reasoning ability of llm. I find your sft model achieve 42.15 on longbenchv2 benchmark, but I check the sft dataset(mainly UltraData-SFT-2605), the user prompts are all very short, mean length less than 500 while user prompts of longbenchv2 are often very long(up to 2M). How does the sft model get the good long context reasoning ability here? From pretraining/mid-training? Or do you use extra long user prompt sft data? Looking forward to your reply!

OpenBMB org

Thank you for your interest. We have not released the relevant long-context data in UltraData-SFT-2605. However, UltraData-RL-2609 does include publicly available long-context RL data. We plan to include more of this data in future releases, so please stay tuned. Additionally, to improve long-context capabilities, we carefully curate a selection of long-form texts in LongDeca stage.

Sign up or log in to comment