TokenNinja-GraphRAG-Inference / data /eval_queries.json
Dhruvpandey1476's picture
Add 30 additional queries (50-query benchmark dataset)
775eb89
Raw History Blame Contribute Delete
21 kB
[
{
"question": "What is the relationship between transformer architecture and attention mechanisms in modern LLMs?",
"ground_truth": "Transformer architecture relies fundamentally on attention mechanisms, specifically self-attention, which allows the model to weigh the importance of different tokens relative to each other. Multi-head attention enables the model to capture multiple types of relationships simultaneously across different representation subspaces."
},
{
"question": "How does BERT differ from GPT in terms of training objectives and use cases?",
"ground_truth": "BERT uses masked language modeling and next sentence prediction for bidirectional context understanding, making it ideal for classification and extraction tasks. GPT uses autoregressive next-token prediction for unidirectional generation, making it better suited for text generation tasks."
},
{
"question": "What are the main causes of hallucination in large language models?",
"ground_truth": "LLM hallucinations stem from training data gaps, overconfident probability distributions, inability to distinguish knowledge boundaries, reward hacking during RLHF, and lack of grounding to external verified knowledge sources."
},
{
"question": "How does retrieval-augmented generation improve factual accuracy compared to parametric knowledge?",
"ground_truth": "RAG improves factual accuracy by grounding model responses in retrieved external documents, reducing reliance on potentially outdated or incorrect parametric knowledge stored in model weights, and providing verifiable source attribution."
},
{
"question": "What is the role of knowledge graphs in reducing LLM token consumption?",
"ground_truth": "Knowledge graphs reduce token consumption by organizing information as structured entity-relationship triples, enabling precise retrieval of only relevant facts rather than large document chunks, resulting in compact structured prompts."
},
{
"question": "How does multi-hop reasoning work in graph-based retrieval systems?",
"ground_truth": "Multi-hop reasoning in graph retrieval follows relationship edges across multiple nodes to connect distant concepts. Starting from seed entities in the query, the system traverses the graph hop by hop, accumulating related entities and relationships that collectively provide richer context than any single document chunk."
},
{
"question": "What is the difference between vector similarity search and graph traversal for RAG?",
"ground_truth": "Vector similarity search retrieves documents by semantic closeness in embedding space, which can miss structural relationships. Graph traversal follows explicit entity-relationship edges, enabling precise multi-hop reasoning and returning only structurally connected facts relevant to the query."
},
{
"question": "How does TigerGraph's GSQL support complex graph analytics?",
"ground_truth": "GSQL is TigerGraph's graph query language that supports multi-hop traversals, pattern matching, accumulator-based aggregations, and parallel computation across the graph, enabling complex analytical queries that would be expensive in traditional databases."
},
{
"question": "What metrics are used to evaluate answer quality in RAG systems?",
"ground_truth": "RAG answer quality is evaluated using BERTScore for semantic similarity, LLM-as-a-Judge for correctness and completeness, ROUGE scores for lexical overlap, faithfulness metrics to check grounding in retrieved context, and human evaluation for overall quality."
},
{
"question": "How does chunking strategy affect RAG performance?",
"ground_truth": "Chunking strategy affects RAG by determining the granularity of retrieved context. Smaller chunks improve precision but may lose surrounding context. Larger chunks provide more context but increase token cost and noise. Overlap between chunks helps preserve cross-boundary information. Semantic chunking by sentence or paragraph boundaries outperforms fixed-size chunking."
},
{
"question": "What are embedding models and how do they enable semantic search?",
"ground_truth": "Embedding models convert text into dense numerical vectors in a high-dimensional space where semantically similar texts are placed close together. Semantic search uses these vectors with similarity metrics like cosine similarity to retrieve documents whose meaning matches the query, even without exact keyword overlap."
},
{
"question": "What is the context window limitation in LLMs and how does RAG address it?",
"ground_truth": "LLMs have fixed context windows limiting how much text they can process at once. RAG addresses this by selectively retrieving only the most relevant document segments, keeping the context within window limits while still drawing on large external knowledge bases."
},
{
"question": "How does FAISS enable efficient similarity search at scale?",
"ground_truth": "FAISS enables efficient similarity search by using approximate nearest neighbor algorithms like IVF (inverted file index) and HNSW (hierarchical navigable small world graphs) that trade small accuracy losses for massive speed gains, allowing billion-scale vector search in milliseconds."
},
{
"question": "What is the significance of BERTScore F1 in evaluating generated text?",
"ground_truth": "BERTScore F1 measures semantic similarity between generated and reference text using contextual BERT embeddings, capturing meaning-level overlap rather than just surface lexical matches. An F1 above 0.88 indicates strong semantic equivalence between the generated answer and the ground truth."
},
{
"question": "How do knowledge graph entity types improve retrieval precision?",
"ground_truth": "Entity type classification (Person, Organization, Concept, Location, Event) allows the retrieval system to filter and prioritize entities relevant to the query type, reducing irrelevant traversal paths and improving the signal-to-noise ratio of retrieved context."
},
{
"question": "What is the LLM-as-a-Judge evaluation methodology?",
"ground_truth": "LLM-as-a-Judge uses a powerful LLM to evaluate generated answers on dimensions like correctness, completeness, relevance, and clarity by comparing them to reference answers or grading them independently. It provides scalable quality assessment without requiring large human annotation teams."
},
{
"question": "How does GraphRAG handle queries that require information from multiple documents?",
"ground_truth": "GraphRAG handles multi-document queries by traversing entity relationship edges that cross document boundaries. Entities mentioned across multiple documents are linked in the graph, allowing the retriever to assemble context from multiple sources through a single graph traversal rather than separate document retrievals."
},
{
"question": "What is the role of confidence scores in knowledge graph relationships?",
"ground_truth": "Confidence scores on knowledge graph relationships indicate the strength of evidence for each entity-entity connection. High-confidence edges (>0.8) represent well-attested relationships while lower scores indicate weaker or inferred connections. Graph retrieval can filter by confidence to return only reliable context."
},
{
"question": "How do prompt engineering techniques reduce token usage?",
"ground_truth": "Prompt engineering reduces token usage through concise instruction phrasing, structured output formats, few-shot examples instead of long explanations, context compression, and providing only task-relevant information rather than exhaustive background."
},
{
"question": "What are the cost implications of token reduction at production scale?",
"ground_truth": "At production scale, a 70% token reduction translates directly to 70% lower LLM API costs, which for systems processing millions of queries daily can represent savings of tens or hundreds of thousands of dollars monthly. It also reduces latency, improves throughput, and decreases the risk of hitting rate limits."
},
{
"question": "What is positional encoding in transformers and why is it necessary?",
"ground_truth": "Positional encoding adds information about token positions to embeddings, enabling transformers to understand word order since self-attention is position-invariant. Common approaches include sinusoidal positional encodings and learned positional embeddings that capture relative and absolute positions in sequences."
},
{
"question": "How does the attention mechanism calculate query, key, and value representations?",
"ground_truth": "The attention mechanism projects input embeddings into query (Q), key (K), and value (V) spaces using learned linear transformations. Attention scores are computed as softmax(Q·K^T/√d), then used to weight and aggregate value vectors, creating context-aware representations that focus on relevant tokens."
},
{
"question": "What is the difference between self-attention and cross-attention?",
"ground_truth": "Self-attention computes attention over the same sequence (Q, K, V all from the same input), allowing tokens to attend to all positions in the input. Cross-attention computes attention from one sequence to another (Q from decoder, K and V from encoder), enabling information transfer between different inputs."
},
{
"question": "How does layer normalization improve transformer training stability?",
"ground_truth": "Layer normalization normalizes activations across feature dimensions to have zero mean and unit variance, reducing internal covariate shift and stabilizing gradient flow during training. It allows higher learning rates, enables training of deeper models, and improves convergence speed compared to batch normalization."
},
{
"question": "What is the role of feed-forward networks in transformer blocks?",
"ground_truth": "Feed-forward networks in transformers consist of two dense layers with a non-linearity (usually ReLU or GELU) between them. They expand the hidden dimension, learn non-linear transformations, and introduce model capacity while maintaining the same output dimension for residual connections."
},
{
"question": "How does scaling relate to transformer model performance?",
"ground_truth": "Transformer performance scales predictably with model size, dataset size, and compute budget following power laws. Larger models with more parameters trained on larger datasets achieve better performance, but with diminishing returns. Scaling laws guide optimal allocation of compute and parameters for a given budget."
},
{
"question": "What are the differences between encoder-only, decoder-only, and encoder-decoder architectures?",
"ground_truth": "Encoder-only models (BERT) use masked language modeling for understanding tasks. Decoder-only models (GPT) use causal masking for generation. Encoder-decoder models (T5) combine both for seq2seq tasks like translation. Each architecture trades off between understanding and generation capabilities."
},
{
"question": "How does instruction tuning improve LLM performance on downstream tasks?",
"ground_truth": "Instruction tuning fine-tunes pre-trained LLMs on diverse instruction-following examples where the model learns to generate responses aligned with specific instructions. This improves zero-shot generalization to new tasks, reduces the need for few-shot examples, and makes models more controllable and safer."
},
{
"question": "What is in-context learning and how does it differ from fine-tuning?",
"ground_truth": "In-context learning (ICL) is where models learn from examples in the prompt without weight updates, enabling fast task adaptation. Fine-tuning updates model weights on task examples, requiring more data and computational resources but potentially achieving better performance for specialized tasks."
},
{
"question": "How does chain-of-thought prompting improve reasoning in LLMs?",
"ground_truth": "Chain-of-thought (CoT) prompting encourages LLMs to generate intermediate reasoning steps before final answers. This improves performance on complex reasoning tasks by decomposing problems into manageable steps, enabling explicit error detection, and allowing the model to correct mistakes during reasoning."
},
{
"question": "What are the main differences between decoding strategies like greedy, beam search, and nucleus sampling?",
"ground_truth": "Greedy decoding selects the highest probability token at each step (fast but may miss better sequences). Beam search explores multiple hypotheses maintaining top-k options. Nucleus (top-p) sampling restricts the vocabulary to the smallest set of tokens with cumulative probability p, balancing diversity and quality."
},
{
"question": "How does reinforcement learning from human feedback improve LLM alignment?",
"ground_truth": "RLHF trains a reward model on human preference pairs, then uses RL (typically PPO) to fine-tune the LLM to maximize reward. This aligns model outputs with human values, reduces harmful content, and improves following user intentions compared to supervised fine-tuning alone."
},
{
"question": "What is the significance of model scale in achieving emergent capabilities?",
"ground_truth": "Emergent capabilities are abilities that appear suddenly as model scale increases but are absent in smaller models. Examples include in-context learning, chain-of-thought reasoning, and instruction following. These emerge due to increased model capacity and diversity of learned representations at larger scales."
},
{
"question": "How do sparse transformers reduce computational complexity?",
"ground_truth": "Sparse transformers use structured sparsity patterns (strided, local, or learned patterns) instead of full attention, reducing complexity from O(n²) to O(n·√n) or O(n). This enables processing longer sequences while maintaining expressive power through strategic attention connections."
},
{
"question": "What is the difference between parametric and non-parametric knowledge in LLMs?",
"ground_truth": "Parametric knowledge is stored in model weights from training data. Non-parametric knowledge comes from external retrieval sources accessed at inference time. RAG combines both: using parametric knowledge for general understanding and non-parametric knowledge for specific, up-to-date facts."
},
{
"question": "How does dense passage retrieval differ from sparse retrieval methods?",
"ground_truth": "Dense retrieval uses learned embeddings for semantic matching, capturing meaning-based similarity. Sparse retrieval uses keyword matching (BM25, TF-IDF), which is more interpretable and efficient but may miss semantic relationships. Modern systems often combine both for complementary strengths."
},
{
"question": "What are hard negatives and why are they important for training dense retrievers?",
"ground_truth": "Hard negatives are non-relevant documents that are semantically similar to queries, making them difficult to distinguish from relevant documents. Training with hard negatives improves retriever robustness and ranking quality by pushing the model to learn finer semantic distinctions beyond surface similarity."
},
{
"question": "How do entity disambiguation techniques improve knowledge graph construction?",
"ground_truth": "Entity disambiguation links mentions of entities to unique identifiers in knowledge bases, resolving ambiguities when the same name refers to different entities. This improves graph quality by preventing spurious entity merging and enabling accurate relationship extraction across mentions of the same entity."
},
{
"question": "What is the relationship between knowledge graph schema design and query performance?",
"ground_truth": "Schema design (entity types, relationship types, properties) affects which queries can be expressed efficiently and how comprehensively. Well-designed schemas enable precise queries through rich type information but require careful balance to avoid over-fragmentation that increases query complexity."
},
{
"question": "How does temporal information in knowledge graphs improve reasoning?",
"ground_truth": "Temporal properties (valid-from, valid-to dates) on entities and relationships allow graphs to reason about historical facts, validity periods, and causal relationships. This enables queries like 'who was the CEO in 2020' and improves factual correctness by acknowledging that facts change over time."
},
{
"question": "What are the tradeoffs between breadth-first and depth-first traversal in graph retrieval?",
"ground_truth": "Breadth-first traversal explores many neighbors at each hop, capturing diverse perspectives but retrieving many entities. Depth-first follows single paths deeply, finding specific connections but potentially missing relevant parallel information. Hybrid strategies or k-hop limits balance both approaches."
},
{
"question": "How does information bottleneck theory relate to token reduction in LLMs?",
"ground_truth": "Information bottleneck theory suggests models learn compressed representations by balancing informativeness and compression. Token reduction achieves compression through graph structure, removing redundant tokens while preserving task-relevant information, improving efficiency without sacrificing quality."
},
{
"question": "What is the role of subgraph summarization in graph-based retrieval?",
"ground_truth": "Subgraph summarization condenses retrieved subgraphs into compact natural language or structured text. This reduces tokens in prompts while preserving essential relationships and context, enabling LLMs to process complex multi-hop connections efficiently while maintaining answer quality."
},
{
"question": "How do graph attention networks combine graph structure with neural attention?",
"ground_truth": "Graph Attention Networks (GATs) use learned attention weights to determine how much each neighbor influences a node's representation. This learns task-specific importance of edges, capturing that not all connections are equally relevant, and improves expressiveness over fixed aggregation schemes."
},
{
"question": "What metrics best capture retriever-reader efficiency in RAG pipelines?",
"ground_truth": "Effective metrics include token count per query, retrieval latency, reader latency, end-to-end latency, cost per query, and quality-adjusted metrics like quality-per-token or quality-per-cost. These capture the full efficiency picture beyond token count alone."
},
{
"question": "How does contrastive learning improve representation quality in dense retrievers?",
"ground_truth": "Contrastive learning minimizes distance between query and relevant passages while maximizing distance from irrelevant passages. This teaches the model to group semantically similar items and separate dissimilar ones, learning robust representations that generalize to new domains and queries."
},
{
"question": "What is the significance of semantic textual similarity datasets for evaluating embeddings?",
"ground_truth": "STS datasets contain sentence pairs with human-assigned similarity scores. Evaluating embeddings on STS provides direct measurement of whether models capture semantic relationships as humans perceive them. Good STS performance correlates with downstream retrieval and ranking effectiveness."
},
{
"question": "How does active learning reduce annotation requirements for training retrievers?",
"ground_truth": "Active learning selects the most informative examples for human annotation rather than random sampling. For retrievers, it identifies hard negatives and ambiguous queries that most improve model performance, reducing annotation cost while maintaining or improving final performance."
},
{
"question": "What are zero-shot and few-shot adaptation techniques for RAG systems?",
"ground_truth": "Zero-shot adaptation uses retrievers and readers trained on general tasks without target task data. Few-shot uses a small number of in-domain examples for prompt engineering or parameter-efficient fine-tuning. These enable rapid deployment to new domains with minimal annotation effort."
},
{
"question": "How do curriculum learning strategies improve RAG system development?",
"ground_truth": "Curriculum learning trains models on progressively harder tasks, starting with simple queries and gradually increasing complexity. For RAG systems, this improves convergence speed and final performance by teaching the model fundamental retrieval patterns before handling complex multi-hop reasoning and rare edge cases."
}
]