Running Reproduction: Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions 🎯 Explore code, traces, and workspace with a collaborative logbook
Running Reproduction: Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions 🎯 Explore code, traces, and workspace with a collaborative logbook
Running Reproduction: DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation 🎯 Explore research agent logs and traces in an interactive workspace
Running Reproduction: DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation 🎯 Explore research agent logs and traces in an interactive workspace
Running Reproduction: Judging What We Cannot Solve 🎯 Explore code logs, traces, and workspace in a web logbook
Running Reproduction: Judging What We Cannot Solve 🎯 Explore code logs, traces, and workspace in a web logbook
Running Reproduction: SEDRAS (WZ-LLM Claims) 🎯 Explore experiment logs, traces, and workspace in a web UI
Running Reproduction: SEDRAS (WZ-LLM Claims) 🎯 Explore experiment logs, traces, and workspace in a web UI
Running Reproduction: Asymmetric Contrastive Objectives for Efficient Phenotypic Screening 🎯 Browse logs, traces, and workspace with a collaborative web logbook
Running Reproduction: Asymmetric Contrastive Objectives for Efficient Phenotypic Screening 🎯 Browse logs, traces, and workspace with a collaborative web logbook
Running Reproduction: Seizure-Semiology-Suite (S3) 🎯 Explore and manage seizure logbook with agent collaboration
Running Reproduction: Seizure-Semiology-Suite (S3) 🎯 Explore and manage seizure logbook with agent collaboration
Running Reproduction: Agentic Framework for Epidemiological Modeling 🎯 Explore model logs and collaborate with an AI agent
Running Reproduction: Agentic Framework for Epidemiological Modeling 🎯 Explore model logs and collaborate with an AI agent
Running Reproduction: ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research 🎯 Explore and manage ORLoopBench benchmark logs
Running Reproduction: ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research 🎯 Explore and manage ORLoopBench benchmark logs
Running Reproduction: Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models 🎯 Explore and manage experiment logs with AI collaboration
Running Reproduction: Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models 🎯 Explore and manage experiment logs with AI collaboration
Running Reproduction: CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research? 🎯 Explore scientific experiment logs, traces, and workspace files
Running Reproduction: CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research? 🎯 Explore scientific experiment logs, traces, and workspace files