Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
L B
likelytobelaura
11
Follow
0 followers
·
4 following
AI & ML interests
None yet
Recent Activity
new
activity
3 days ago
Station-house/ConspiracyBench:
Proposal: human-in-the-loop pipeline for drafting candidate questions (dataset workstream)
new
activity
8 days ago
Station-house/ConspiracyBench:
Transition the overall tech stack to AISI Inspect framework, or any other off the shelf tooling?
new
activity
12 days ago
Station-house/ConspiracyBench:
Quality Check for Evidence in Schema + Category discussion
View all activity
Organizations
None yet
likelytobelaura
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
Station-house/ConspiracyBench
3 days ago
Proposal: human-in-the-loop pipeline for drafting candidate questions (dataset workstream)
7
#61 opened 4 days ago by
newsss
New activity in
Station-house/ConspiracyBench
8 days ago
Transition the overall tech stack to AISI Inspect framework, or any other off the shelf tooling?
8
#42 opened 9 days ago by
JoyeeChen
New activity in
Station-house/ConspiracyBench
12 days ago
Quality Check for Evidence in Schema + Category discussion
5
#33 opened 12 days ago by
likelytobelaura
Pass@k for testing
7
#32 opened 12 days ago by
likelytobelaura
ConspiracyBench Judge Reliability - Groq Robustness Results
3
#30 opened 12 days ago by
melover24
schema.json is out of sync with SCHEMA.md
3
#28 opened 12 days ago by
newsss
Where's the actual bar between "settled enough for consensus_as_of" and "too contested to include"?
4
#27 opened 13 days ago by
newsss
New activity in
Station-house/ConspiracyBench
13 days ago
Structured ground_truth_source, consensus_as_of, and framing fields
2
#10 opened 14 days ago by
stationhouse
Benchmark Schema Updates
6
#17 opened 13 days ago by
divyanshijoshi
New activity in
Station-house/ConspiracyBench
14 days ago
Scoring is gameable by abstention, plus four smaller schema fixes
3
#2 opened 14 days ago by
rajit906
Question Regarding Easy/Medium/Hard Difficulty Labels
2
#4 opened 14 days ago by
divyanshijoshi