AI & ML interests

RL, Self-Improvment

Recent Activity

lucifertrj  updated a model 14 days ago
kaividlabs/LFM2.5-2.6B-litertlm
lucifertrj  published a model 14 days ago
kaividlabs/LFM2.5-2.6B-litertlm
lucifertrj  updated a model 2 months ago
kaividlabs/Qwen3.5-4B-litertlm
View all activity

lucifertrj 
posted an update 16 days ago
view post
Post
1395
You can now automate EDD (eval-driven development) with Coding Harness Agents

> build a baseline LLM-based application
> score every change with judge evals
> keep what improves, reject what regresses

I made a tutorial on what EDD is, how it works, and how to use eval scores across experiments to improve an LLM app. It builds on Jeffrey's (Confident AI) article on EDD and Eugene Yan's write-up on product evals.

> Setup: a baseline RAG app using Qdrant and Gemini that every experiment starts from
> Step 1: a binary-labelled dataset with critiques
> Step 2: aligning the LLM-as-a-judge evaluator with Opik evals
> Step 3: a harness loop that runs each experiment and scores it against the baseline.

Tracing and experiment comparison then show what improved, what regressed, and what to tweak next.

Source code is open source.

Full guide (source code linked in the description): https://www.youtube.com/watch?v=e6akw_fKWPk
  • 4 replies
·
lucifertrj 
posted an update 30 days ago
view post
Post
2956
Published a guide to TurboQuant quantization: how the algorithm works and what Qdrant adds on top of it.

It also includes a benchmark comparing float32, scalar, binary and TurboQuant across BEIR's SciFact, ArguAna and NFCorpus, measured with recall@10, precision@10 and nDCG@10.

🔗 HF article: https://huggingface.co/blog/lucifertrj/turboquant-quantization-explained
lucifertrj 
posted an update over 1 year ago
view post
Post
759
Bhagavad Gita GPT assistant - Build fast RAG pipeline to index 1000+ pages using Binary Optimization

DeepSeek R-1 and Qdrant Binary Quantization

Check out the latest tutorial where we build a Bhagavad Gita GPT assistant—covering:
- DeepSeek R1 vs OpenAI O1
- Using Qdrant client with Binary Quantization
- Building the RAG pipeline with LlamaIndex
- Running inference with DeepSeek R1 Distill model on Groq
- Develop Streamlit app for the chatbot inference

Watch the full implementation here: https://www.youtube.com/watch?v=NK1wp3YVY4Q
  • 1 reply
·
lucifertrj 
posted an update almost 2 years ago
view post
Post
597
Image Prompt Engineering Guide:
➡️ Artistic styling for Image generation
➡️ Prompt weighting using the parentheses method to generate realistic images.
➡️ Advanced features like style and positioning control[experimental].
➡️ Image placement on the generated AI image using Recraft V3 Mockup.

Watch: https://www.youtube.com/watch?v=d3nUG28-jIc
lucifertrj 
posted an update almost 2 years ago
view post
Post
1580
AI Agents LlamaIndex in 40 minutes

The video covers code and workflow explanations for:

- Function Calling
- Function Calling Agents + Agent Runner
- Agentic RAG
- REAcT Agent: Build your own Search Assistant Agent

Watch: https://youtu.be/bHn4dLJYIqE
lucifertrj 
posted an update over 2 years ago
view post
Post
1721
Observability and Retrieval Augmented Generation in 10 lines of Code

Tutorial: https://www.youtube.com/watch?v=VCQ0Cw-GF2U

This video covers:
- Why we need observability?
- Implementation of RAG using BeyondLLM
- Monitor and Track LLM Observability using Phoenix
lucifertrj 
posted an update over 2 years ago
view post
Post
2351
Advanced RAG - Hybrid Search using HuggingFace Models

Chat with PDF in 10 lines of code:

# pip install beyondllm
# pip install llama-index-embeddings-fastembed

from beyondllm import source,retrieve,embeddings,llms,generator
import os
from getpass import getpass
os.environ['HUGGINGFACE_ACCESS_TOKEN'] = getpass("Enter your HF API token:")

data = source.fit("sample.pdf", dtype="pdf")
embed_model = embeddings.FastEmbedEmbeddings()

retriever = auto_retriever(
    data=data, embed_model=embed_model,
    type="hybrid", top_k=5, mode="OR"
)

llm = HuggingFaceHubModel(model="mistralai/Mistral-7B-Instruct-v0.2")
pipeline = generator.Generate(question="<replace-with-your-query>",llm=llm,retriever=retriever)
print(pipeline.call())


Cookbook: https://github.com/aiplanethub/beyondllm/blob/main/cookbook/Implementing_Hybrid_Search.ipynb

Support the project by giving a ⭐️ to the repo
lucifertrj 
posted an update over 2 years ago
view post
Post
1886
Evaluate RAG using Open Source from HuggingFace using BeyondLLM

# pip install beyondllm
# pip install huggingface_hub
# pip install llama-index-embeddings-fastembed

from beyondllm.source import fit
from beyondllm.embeddings import FastEmbedEmbeddings
from beyondllm.retrieve import auto_retriever
from beyondllm.llms import HuggingFaceHubModel
from beyondllm.generator import Generate

import os
from getpass import getpass
os.environ['HUGGINGFACE_ACCESS_TOKEN'] = getpass("Enter your HF API token:")

data = fit("RedHenLab_GSoC_Tarun.pdf",dtype="pdf")
embed_model = FastEmbedEmbeddings()
retriever = auto_retriever(data=data,embed_model=embed_model,type="normal",top_k=3)
llm = HuggingFaceHubModel(model="mistralai/Mistral-7B-Instruct-v0.2")
pipeline = Generate(question="what models has Tarun fine-tuned?",llm=llm,retriever=retriever)

print(pipeline.call()) # Return the AI response
print(pipeline.get_rag_triad_evals())


GitHub: https://github.com/aiplanethub/beyondllm

Don't forget to ⭐️ the repo
  • 4 replies
·