-
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
Paper β’ 2605.25874 β’ Published β’ 81 -
meituan-longcat/WBench
Benchmark β’ Updated β’ 867 β’ 2.23k β’ 25 -
meituan-longcat/WBench-weights
Other β’ Updated β’ 11 -
meituan-longcat/WBench-examples
Viewer β’ Updated β’ 447 β’ 362 β’ 4
AI & ML interests
None defined yet.
Recent Activity
Papers
DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
π¦π½πͺβ¨β¨
γNewγ LongCat-DeepResearch
LongCat-DeepResearch combines a LongCat model enhanced for deep research with a research harness for open-ended investigation and evidence-grounded report generation. The system outperforms the evaluated deep-research offerings from ChatGPT, Claude, and Gemini across DeepResearchBench, DeepResearchBench II, and ResearchRubrics. These model improvements are intended to be incorporated into the next general-purpose LongCat model release.
Welcome to LongCat Home
Open foundation models for long-context reasoning, agentic coding, and scalable AI systems. Home of LongCat-2.0.
LongCat is a large language model family built by Meituan. We're working on making AI more useful in the physical world β one small step at a time. We continuously release LLMs, multimodal models, and other AI projects here. Feel free to visit LongCat and enjoy our latest models!
Latest innovations
- LongCat-DeepResearch: Technical Report | GitHub | Blog
- LongCat-2.0: Technical Report/Demo | Github |
- LongCat-Video-Avatar-1.5: Technical Report | Project Page
- LongCat-Next: Technical Report | Github | Project Page | Demo
- LongCat-AudioDiT: Technical Report | Github | Demo
- LongCat-Flash-Prover: Technical Report | Github
- LongCat-Video: Technical Report | Github | Project Page
- LongCat Video-avatar: Technical Report | Github | Project Page
- LongCat-Flash-Omni: Technical Report | Github
- LongCat-Image: Technical Report | Github
- LongCat-Flash-Chat: Technical Report | Github
Explore More
-
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
Paper β’ 2605.25874 β’ Published β’ 81 -
meituan-longcat/WBench
Benchmark β’ Updated β’ 867 β’ 2.23k β’ 25 -
meituan-longcat/WBench-weights
Other β’ Updated β’ 11 -
meituan-longcat/WBench-examples
Viewer β’ Updated β’ 447 β’ 362 β’ 4