File size: 14,600 Bytes
e0f0f69
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ad0f40
e0f0f69
9ad0f40
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e0f0f69
 
 
 
 
549df35
9ad0f40
 
 
 
 
 
 
 
cc32148
2c53aea
7c6c442
9ad0f40
 
 
 
 
e0f0f69
 
 
 
 
cc32148
 
9ad0f40
 
 
cef3ff2
549df35
2c53aea
9ad0f40
764e308
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ad0f40
 
 
e0f0f69
 
 
9ad0f40
 
 
 
 
 
 
 
 
 
 
 
 
764e308
 
9ad0f40
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c5714b4
eab9341
 
 
c5714b4
 
 
eab9341
c5714b4
 
 
 
 
6e4b914
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ad0f40
6e4b914
764e308
1b90ba6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ad0f40
e0f0f69
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <meta name="viewport" content="width=device-width, initial-scale=1.0">
    <title>A Graph RAG study - Experimental Setup</title>
    <style>
        body {
            font-family: Arial, sans-serif;
            line-height: 1.6;
            margin: 20px;
        }
        h1, h2 {
            color: #333;
        }
        h2 {
            margin-top: 30px;
        }
        ul {
            list-style-type: disc;
            margin-left: 20px;
        }
        p {
            margin-bottom: 15px;
        }

        table{
            border-collapse: collapse;
            width: 95%;
            border: 2px solid #2c3e50;
        }
        
        tr{
            border-bottom: 2px solid #b60e0e;
        }

        td{
            width: 15%;
            vertical-align: top;
            border: 2px solid #3498db;
        }
    </style>
</head>
<body>

    <h1>An experimental setup for testing Hybrid RAG (Vector and Graph RAG)</h1>

    <p>
        Continue to read this article if you want to experiment with Graph RAG.
        We used UV for python package management.
        Following is our setup. 
      </p>  
        <table border="1">
            <thead>
            <tr>
                <th>Vector DB</th>
                <th>Observability Tool</th>
                <th>LLM of choice</th>
                <th>Embedding Model</th>
                <th>Graph DB</th>
                <th>Orchestration Framework</th>
            </tr>
            </thead>
            <tbody>
            <tr>
            <td>
                <p>
                Vector DB setup: (not using the latest chromadb, we want to experiment with datapipes, hence this version)
                <br>uv pip list | grep chroma<br>
                Using Python 3.12.12 environment at: .<br>
                chroma-hnswlib                               0.7.6<br>
                chromadb                                     0.5.23<br>
                
                    export CHROMA_OTEL_COLLECTION_ENDPOINT="http://{host}:4317"
                    export CHROMA_OTEL_SERVICE_NAME="chroma-server-local"
                    export CHROMA_OTEL_GRANULARITY="all"

                    HOST={host}
                    PORT={port}
                    DB="{database path}"
                    chroma run --host=$HOST --port=$PORT --path=$DB
               
                </p>
                
            </td>

            <td>
                 <p>
                    Open Telemetry: Phoenix (This part is optional, but good to have a dashboard for observability)
                <br>uv pip list | grep phoenix<br>
                arize-phoenix                                19.4.0<br>
                arize-phoenix-client                         2.13.0<br>
                arize-phoenix-evals                          3.2.0<br>
                arize-phoenix-otel                           0.16.1<br>
                 </p> 
                 <p>
                    Running phoenix server: phoenix serve
                    Once you have phoenix observability server running and chroma db is sending telemetry data to the server.
                    You can view the results in the dashboard.
                    We will use Phoenix as LLM observability tool. We can get llama index to broadcast execution data to Phoenix<br>
                    
                </p> 
               
            </td>

            <td>
                    LLM: google_gemma-3-1b-it-Q6_K.llamafile (You can use any llm, but we wanted to use open_ai_like)
            </td>
               
            <td>
                    <p>
                        Embedding Model: We used Ollama to serve the model
                <br>ollama list
                <br>NAME                     Dimensions             
                <br>nomic-embed-text:latest  768         
                <br>all-minilm:l6-v2         384  
                    </p>
            </td>
            
            <td>
                    <p>
                    Graph Database (Inmemory/Kuzu)
                <br>uv pip list | grep kuzu
                <br>kuzu                                     0.11.3
                <br>llama-index-graph-stores-kuzu            0.9.1
                </p>
            </td>

            <td>
                    Orchestration Framework: Llamaindex
            </td>

            </tr>
        </tbody>

        </table>
       
   
    
    <pre>
                    +----------------------+

                    |  Input Data Sources  |
                    +---------+------------+
                                |
                    +-----------v-----------------+

                    | Data Ingestion & Processing |
                    +-----+------------+----------+
                          |            |
                          |            |
       +------------------+            +------------------+

       |                                                  |
+------v-------------------+                       +------v-------------------+

| Knowledge Graph Const.   |                       | Text Chunking & Embed.   |
| (Extract Entities/Rels)  |                       | (Generate Vectors)       |
+------+-------------------+                       +------+-------------------+

       |                                                  |
+------v-------------------+                       +------v-------------------+

| Graph Database           |                       | Vector Database          |
| (Kuzu)                   |                       | (Chroma DB)              |
+------+-------------------+                       +------+-------------------+

       |                                                  |
       |  Structural Retrieval                            |  Semantic Search
       |  (Multi-hop Traversal)                           |  (Cosine/Euclidean)
       +------------------+            +------------------+

                          |            |
                    +-----v------------v-----+

                    | Hybrid Query Engine    |
                    |      & Retriever       |
                    +------------+-----------+
                                 |
                    +------------v-----------+

                    |  Context Fusion & LLM  |
                    |  (Re-ranking & Synth)  |
                    +------------+-----------+
                                 |
                    +------------v-----------+

                    |   Generated Response   |
                    +------------------------+
    </pre>
    <h3>Practice time</h3>
    <p>
        <h3>Data Ingestion pipeline</h3> 
        Transforming chunks of text into Vectors (384, 768 dimensions etc), provides a way to capture semantic meaning.
        <br>This enables search by similarity in meaning, not exact search.
        <br> Input Document(s)/Images/Audio/Video --> Appropriately Chunked (chopped to a right size) --> Embedded (Vectorized) --> Persisted to a Vector Database
        <ul>
          <li>langchain-textsplitters library provides several useful specialized chunkers like RecursiveCharacterTextSplitter, SemanticSplitter etc to chunk</li>
          <li>Chunking size(500 to 1000 tokens), chunking overlap (10% to 20%) impacts the quality of search</li>
          <li>Choosing the right emedding model like all-minilm, nomic-embed etc</li>
          <li>Choosing the right vector data store, chromadb, pinecone, milvus etc</li>
          <li>Setting up an evaluation pipeline to validate the data integrity</li>
        </ul>
    </p>
    <p>
        <b>Understanding some basic chunking</b><br>
        When a large text or a large document needs to be avilable for RAG (Retrieval Augmented Generation), we chop them into smaller chunks.
        A chunk is a small piece of information. Often, a good vectorization is to optimize chunking, adjusting the sliding window and choosing the right kind of embedding model.
        We will deep dive into this topic in subsequent practice sessions.
        One good library that provides multiple splitters (chunkers) is langchain. Get it installed using uv.
        We are using <i>langchain-text-splitters == 1.1.2</i><br>
        Here is a very simple example to clearly see how it works if we used a token splitter.
        <pre>
            from langchain_text_splitters import TokenTextSplitter
            text_splitter = TokenTextSplitter(
                chunk_size=20,  # tokens per each chunk
                chunk_overlap=5 # shared tokens in each chunk
            )
            raw_text = ("Choose any paragraph of your choice")
            chunks = text_splitter.create_documents([raw_text]) # you may want to check the size by using len(chunks)
            for i, chunk in enumerate(chunks):
                print(f"\n--- Chunk {i+1} ---")
                print(chunk.page_content)
        </pre>

    </p>
    <p>
        <i>Going little further, trying to understand how search would work on this</i>
        <pre>
            import json
            from langchain_text_splitters import RecursiveJsonSplitter

            # 1. Define your tabular data as a list of row dictionaries
            tabular_data = [
                {"Employee ID": 101, "Name": "Adi", "Department": "Engineering"},
                {"Employee ID": 102, "Name": "Bhimsen", "Department": "Marketing"},
                {"Employee ID": 103, "Name": "Chota", "Department": "Sales"},
                {"Employee ID": 104, "Name": "No Name"},
            ]
            # 2. We will use a RecursiveJsonSplitter
            splitter = RecursiveJsonSplitter(max_chunk_size=100) 

            # 3. Split the data into LangChain Document objects
            # This formats the JSON and converts it to search-ready text chunks
            documents = splitter.create_documents(texts=[tabular_data], convert_lists=True)

            # 4. Like prevous example, we can check the chunks that are created.
            text_chunks = []
            for i, doc in enumerate(documents):
                print(f"--- Chunk {i+1} ---")
                print(doc.page_content)
                text_chunks.append(doc.page_content)

            # 5. #now we will do some search operations on the chunks
            from sentence_transformers import SentenceTransformer, util
            model = SentenceTransformer("all-MiniLM-L6-v2") # you may choose other embedding models
            corpus_embeddings = model.encode(text_chunks, convert_to_tensor=True)

            # 6. Now you can truly understand, how thise works when we do a search
            # Calculate similarity strictly in-memory via tensor dot-product. There are two questions.
            # Play with the top_k value for finding the matches
            query_embedding = model.encode("who works in engineering deparment?", convert_to_tensor=True)
            #query_embedding = model.encode("who do not work in marketing department?", convert_to_tensor=True)
            #hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=1)
            hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=5)

            # 7. Check the results
            for rank, hit in enumerate(hits[0]):
                chunk_index = hit['corpus_id']
                score = hit['score']
                print(f"Rank {rank+1} (Score: {score:.4f}): {text_chunks[chunk_index]}")
        </pre>
    </p>
    
    <h3>Detailed article coming soon ..</h3>
     
    <!-- 
    <table >
        <tr>
            <th>The time telling myth</th>
            <th>Why bloom in the evening</th>
            <th>What type of pollinators are attracted</th>
            
        </tr>
        <tr>
          <td>
            How Exactly Does the Four O’Clock Plant Tell the Time? The four o’clock plant, and many other plants are able to “tell the time” because of pigments such as phytochromes that are able to detect the amount of daylight. 
          </td>
            <td>
               The four o’clock plant particularly opens at 4 PM to avoid competition with most other flowers, which open in mornings to attract pollinators such as bees. 
                At 4 PM, these pollinators are less active and nocturnal pollinators like moths, hummingbirds, and butterflies are active instead. 

            </td>
            <td>
                Four o’clock plants have a distinct shape, and are attracted by a specific audience of pollinators, more specifically those with long tongues that are able to consume the nectar. 
            </td>
            
        </tr>
    </table>

    <h4>Starting your own garden</h4>

    <table style="border-collapse: collapse;">
        <tr>
            <th>Initial Stage</th>
            <th>Small plant stage</th>
            <th>Intermediate stage</th>
            <th>Bloom Time</th>
        </tr>
        <tr >
            <td>
             Soaking seeds in water before planting them is helpful and can increase germination rates, since the seeds are typically quite hard.
             <img src="images/Four_O_Clock.jpeg" height="200" width="200"/>
            </td>
            <td>
              The plants require regular watering (though overwatering can be detrimental), and in some cases adding mulch around them can be helpful to retain moisture.   
              <img src="images/Four_O_Clock_Garden.jpeg" height="200" width="200"/>
            </td>
            <td>
              These plants typically bloom in the summer all the way until fall, and thrive in warmer weather. 
                <img src="images/Mirabilis_Closeup.jpeg" height="200" width="200"/>
            </td>
            <td>
                These plants die in the winter, and come back in the following spring, as the roots underneath the soil stay alive and will push for new growth once temperature increases.
                <img src="images/Mirabilis_Full_Circl.jpeg" height="200" width="200"/>
            </td>
        </tr>
    </table>
   <h3>Conclusion</h3>
    <p>Adding these wonderful environment friendly plants to your garden can be very rewarding experience that offers a sustainable and fulfilling way to connect with nature.  
        It's a simple yet powerful way to reduce your environmental impact, save money, and enjoy this beautiful colorful plant in your garden.    
        You can save the seeds, which are produced in abundance and can be shared with friends and family. Start your journey today!</p>
    -->

</body>
</html>