Spaces:
Running
Running
File size: 14,600 Bytes
e0f0f69 9ad0f40 e0f0f69 9ad0f40 e0f0f69 549df35 9ad0f40 cc32148 2c53aea 7c6c442 9ad0f40 e0f0f69 cc32148 9ad0f40 cef3ff2 549df35 2c53aea 9ad0f40 764e308 9ad0f40 e0f0f69 9ad0f40 764e308 9ad0f40 c5714b4 eab9341 c5714b4 eab9341 c5714b4 6e4b914 9ad0f40 6e4b914 764e308 1b90ba6 9ad0f40 e0f0f69 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 | <!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>A Graph RAG study - Experimental Setup</title>
<style>
body {
font-family: Arial, sans-serif;
line-height: 1.6;
margin: 20px;
}
h1, h2 {
color: #333;
}
h2 {
margin-top: 30px;
}
ul {
list-style-type: disc;
margin-left: 20px;
}
p {
margin-bottom: 15px;
}
table{
border-collapse: collapse;
width: 95%;
border: 2px solid #2c3e50;
}
tr{
border-bottom: 2px solid #b60e0e;
}
td{
width: 15%;
vertical-align: top;
border: 2px solid #3498db;
}
</style>
</head>
<body>
<h1>An experimental setup for testing Hybrid RAG (Vector and Graph RAG)</h1>
<p>
Continue to read this article if you want to experiment with Graph RAG.
We used UV for python package management.
Following is our setup.
</p>
<table border="1">
<thead>
<tr>
<th>Vector DB</th>
<th>Observability Tool</th>
<th>LLM of choice</th>
<th>Embedding Model</th>
<th>Graph DB</th>
<th>Orchestration Framework</th>
</tr>
</thead>
<tbody>
<tr>
<td>
<p>
Vector DB setup: (not using the latest chromadb, we want to experiment with datapipes, hence this version)
<br>uv pip list | grep chroma<br>
Using Python 3.12.12 environment at: .<br>
chroma-hnswlib 0.7.6<br>
chromadb 0.5.23<br>
export CHROMA_OTEL_COLLECTION_ENDPOINT="http://{host}:4317"
export CHROMA_OTEL_SERVICE_NAME="chroma-server-local"
export CHROMA_OTEL_GRANULARITY="all"
HOST={host}
PORT={port}
DB="{database path}"
chroma run --host=$HOST --port=$PORT --path=$DB
</p>
</td>
<td>
<p>
Open Telemetry: Phoenix (This part is optional, but good to have a dashboard for observability)
<br>uv pip list | grep phoenix<br>
arize-phoenix 19.4.0<br>
arize-phoenix-client 2.13.0<br>
arize-phoenix-evals 3.2.0<br>
arize-phoenix-otel 0.16.1<br>
</p>
<p>
Running phoenix server: phoenix serve
Once you have phoenix observability server running and chroma db is sending telemetry data to the server.
You can view the results in the dashboard.
We will use Phoenix as LLM observability tool. We can get llama index to broadcast execution data to Phoenix<br>
</p>
</td>
<td>
LLM: google_gemma-3-1b-it-Q6_K.llamafile (You can use any llm, but we wanted to use open_ai_like)
</td>
<td>
<p>
Embedding Model: We used Ollama to serve the model
<br>ollama list
<br>NAME Dimensions
<br>nomic-embed-text:latest 768
<br>all-minilm:l6-v2 384
</p>
</td>
<td>
<p>
Graph Database (Inmemory/Kuzu)
<br>uv pip list | grep kuzu
<br>kuzu 0.11.3
<br>llama-index-graph-stores-kuzu 0.9.1
</p>
</td>
<td>
Orchestration Framework: Llamaindex
</td>
</tr>
</tbody>
</table>
<pre>
+----------------------+
| Input Data Sources |
+---------+------------+
|
+-----------v-----------------+
| Data Ingestion & Processing |
+-----+------------+----------+
| |
| |
+------------------+ +------------------+
| |
+------v-------------------+ +------v-------------------+
| Knowledge Graph Const. | | Text Chunking & Embed. |
| (Extract Entities/Rels) | | (Generate Vectors) |
+------+-------------------+ +------+-------------------+
| |
+------v-------------------+ +------v-------------------+
| Graph Database | | Vector Database |
| (Kuzu) | | (Chroma DB) |
+------+-------------------+ +------+-------------------+
| |
| Structural Retrieval | Semantic Search
| (Multi-hop Traversal) | (Cosine/Euclidean)
+------------------+ +------------------+
| |
+-----v------------v-----+
| Hybrid Query Engine |
| & Retriever |
+------------+-----------+
|
+------------v-----------+
| Context Fusion & LLM |
| (Re-ranking & Synth) |
+------------+-----------+
|
+------------v-----------+
| Generated Response |
+------------------------+
</pre>
<h3>Practice time</h3>
<p>
<h3>Data Ingestion pipeline</h3>
Transforming chunks of text into Vectors (384, 768 dimensions etc), provides a way to capture semantic meaning.
<br>This enables search by similarity in meaning, not exact search.
<br> Input Document(s)/Images/Audio/Video --> Appropriately Chunked (chopped to a right size) --> Embedded (Vectorized) --> Persisted to a Vector Database
<ul>
<li>langchain-textsplitters library provides several useful specialized chunkers like RecursiveCharacterTextSplitter, SemanticSplitter etc to chunk</li>
<li>Chunking size(500 to 1000 tokens), chunking overlap (10% to 20%) impacts the quality of search</li>
<li>Choosing the right emedding model like all-minilm, nomic-embed etc</li>
<li>Choosing the right vector data store, chromadb, pinecone, milvus etc</li>
<li>Setting up an evaluation pipeline to validate the data integrity</li>
</ul>
</p>
<p>
<b>Understanding some basic chunking</b><br>
When a large text or a large document needs to be avilable for RAG (Retrieval Augmented Generation), we chop them into smaller chunks.
A chunk is a small piece of information. Often, a good vectorization is to optimize chunking, adjusting the sliding window and choosing the right kind of embedding model.
We will deep dive into this topic in subsequent practice sessions.
One good library that provides multiple splitters (chunkers) is langchain. Get it installed using uv.
We are using <i>langchain-text-splitters == 1.1.2</i><br>
Here is a very simple example to clearly see how it works if we used a token splitter.
<pre>
from langchain_text_splitters import TokenTextSplitter
text_splitter = TokenTextSplitter(
chunk_size=20, # tokens per each chunk
chunk_overlap=5 # shared tokens in each chunk
)
raw_text = ("Choose any paragraph of your choice")
chunks = text_splitter.create_documents([raw_text]) # you may want to check the size by using len(chunks)
for i, chunk in enumerate(chunks):
print(f"\n--- Chunk {i+1} ---")
print(chunk.page_content)
</pre>
</p>
<p>
<i>Going little further, trying to understand how search would work on this</i>
<pre>
import json
from langchain_text_splitters import RecursiveJsonSplitter
# 1. Define your tabular data as a list of row dictionaries
tabular_data = [
{"Employee ID": 101, "Name": "Adi", "Department": "Engineering"},
{"Employee ID": 102, "Name": "Bhimsen", "Department": "Marketing"},
{"Employee ID": 103, "Name": "Chota", "Department": "Sales"},
{"Employee ID": 104, "Name": "No Name"},
]
# 2. We will use a RecursiveJsonSplitter
splitter = RecursiveJsonSplitter(max_chunk_size=100)
# 3. Split the data into LangChain Document objects
# This formats the JSON and converts it to search-ready text chunks
documents = splitter.create_documents(texts=[tabular_data], convert_lists=True)
# 4. Like prevous example, we can check the chunks that are created.
text_chunks = []
for i, doc in enumerate(documents):
print(f"--- Chunk {i+1} ---")
print(doc.page_content)
text_chunks.append(doc.page_content)
# 5. #now we will do some search operations on the chunks
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer("all-MiniLM-L6-v2") # you may choose other embedding models
corpus_embeddings = model.encode(text_chunks, convert_to_tensor=True)
# 6. Now you can truly understand, how thise works when we do a search
# Calculate similarity strictly in-memory via tensor dot-product. There are two questions.
# Play with the top_k value for finding the matches
query_embedding = model.encode("who works in engineering deparment?", convert_to_tensor=True)
#query_embedding = model.encode("who do not work in marketing department?", convert_to_tensor=True)
#hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=1)
hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=5)
# 7. Check the results
for rank, hit in enumerate(hits[0]):
chunk_index = hit['corpus_id']
score = hit['score']
print(f"Rank {rank+1} (Score: {score:.4f}): {text_chunks[chunk_index]}")
</pre>
</p>
<h3>Detailed article coming soon ..</h3>
<!--
<table >
<tr>
<th>The time telling myth</th>
<th>Why bloom in the evening</th>
<th>What type of pollinators are attracted</th>
</tr>
<tr>
<td>
How Exactly Does the Four O’Clock Plant Tell the Time? The four o’clock plant, and many other plants are able to “tell the time” because of pigments such as phytochromes that are able to detect the amount of daylight.
</td>
<td>
The four o’clock plant particularly opens at 4 PM to avoid competition with most other flowers, which open in mornings to attract pollinators such as bees.
At 4 PM, these pollinators are less active and nocturnal pollinators like moths, hummingbirds, and butterflies are active instead.
</td>
<td>
Four o’clock plants have a distinct shape, and are attracted by a specific audience of pollinators, more specifically those with long tongues that are able to consume the nectar.
</td>
</tr>
</table>
<h4>Starting your own garden</h4>
<table style="border-collapse: collapse;">
<tr>
<th>Initial Stage</th>
<th>Small plant stage</th>
<th>Intermediate stage</th>
<th>Bloom Time</th>
</tr>
<tr >
<td>
Soaking seeds in water before planting them is helpful and can increase germination rates, since the seeds are typically quite hard.
<img src="images/Four_O_Clock.jpeg" height="200" width="200"/>
</td>
<td>
The plants require regular watering (though overwatering can be detrimental), and in some cases adding mulch around them can be helpful to retain moisture.
<img src="images/Four_O_Clock_Garden.jpeg" height="200" width="200"/>
</td>
<td>
These plants typically bloom in the summer all the way until fall, and thrive in warmer weather.
<img src="images/Mirabilis_Closeup.jpeg" height="200" width="200"/>
</td>
<td>
These plants die in the winter, and come back in the following spring, as the roots underneath the soil stay alive and will push for new growth once temperature increases.
<img src="images/Mirabilis_Full_Circl.jpeg" height="200" width="200"/>
</td>
</tr>
</table>
<h3>Conclusion</h3>
<p>Adding these wonderful environment friendly plants to your garden can be very rewarding experience that offers a sustainable and fulfilling way to connect with nature.
It's a simple yet powerful way to reduce your environmental impact, save money, and enjoy this beautiful colorful plant in your garden.
You can save the seeds, which are produced in abundance and can be shared with friends and family. Start your journey today!</p>
-->
</body>
</html> |