Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Michal K's picture

Michal K

krawcowo
3
Maggio33's profile picture NolaKara's profile picture kacperwikiel's profile picture
·

AI & ML interests

None yet

Recent Activity

repliedto lucifertrj's post 2 days ago
You can now automate EDD (eval-driven development) with Coding Harness Agents > build a baseline LLM-based application > score every change with judge evals > keep what improves, reject what regresses I made a tutorial on what EDD is, how it works, and how to use eval scores across experiments to improve an LLM app. It builds on Jeffrey's (Confident AI) article on EDD and Eugene Yan's write-up on product evals. > Setup: a baseline RAG app using Qdrant and Gemini that every experiment starts from > Step 1: a binary-labelled dataset with critiques > Step 2: aligning the LLM-as-a-judge evaluator with Opik evals > Step 3: a harness loop that runs each experiment and scores it against the baseline. Tracing and experiment comparison then show what improved, what regressed, and what to tweak next. Source code is open source. Full guide (source code linked in the description): https://www.youtube.com/watch?v=e6akw_fKWPk
liked a dataset 6 days ago
PiotrSty/sejm-committee-transcripts
liked a model 14 days ago
kacperwikiel/RysOCR
View all activity

Organizations

Fabryka AI's profile picture

models 0

None public yet

datasets 0

None public yet
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs