tahamajs's picture
|
download
raw
6.71 kB

Getting Started with CA10 Knowledge Graphs

๐Ÿš€ Quick Start Guide

This guide will help you get started with the CA10 Knowledge Graphs project in just a few minutes.

Prerequisites

Before you begin, ensure you have:

  • Python 3.8 or higher installed
  • pip package manager
  • Git (for cloning the repository)
  • Optional: CUDA-capable GPU for faster training

Step 1: Installation

Clone the Repository

git clone <repository-url>
cd CA10_knowledge_graphs

Install Dependencies

pip install -r config/requirements.txt

This will install all required packages including:

  • PyTorch and related libraries
  • NetworkX for graph operations
  • LangChain and LangGraph for reasoning
  • FastAPI for API server
  • Visualization libraries

Step 2: Configuration

Set Up Environment Variables (Optional)

If you want to use advanced features like Gemini API or RAG systems:

# Copy the example configuration
cp config/env_example.txt .env

# Edit .env with your API keys
nano .env  # or use your preferred editor

Add your API keys:

GOOGLE_API_KEY=your_gemini_api_key_here
OPENAI_API_KEY=your_openai_key_here
ANTHROPIC_API_KEY=your_anthropic_key_here

Step 3: Run the Project

Basic Execution

cd scripts
./run.sh

This will:

  1. Create necessary directories
  2. Install any missing dependencies
  3. Run the main knowledge graph analysis
  4. Generate visualizations and reports
  5. Save results to data/ folder

Manual Execution

If you prefer to run components individually:

# Run basic knowledge graph creation
python src/main_execution.py

# Run advanced features with AI integration
python integrations/main_execution_advanced.py

# Start the API server
python src/start_api.py

Step 4: Explore the Results

After execution, check the following directories:

Visualizations

ls data/visualizations/

You'll find:

  • basic_knowledge_graph.png - Basic KG visualization
  • advanced_scientific_network.png - Scientific network analysis
  • transe_embeddings.png - TransE model embeddings
  • distmult_embeddings.png - DistMult model embeddings
  • complex_embeddings.png - ComplEx model embeddings

Results

ls data/results/

Contains JSON files with:

  • Scientific knowledge graph analysis
  • Network metrics and statistics
  • Influence scores and rankings

Logs

tail -f data/logs/*.log

View execution logs and debugging information.

Step 5: Try the Notebooks

Explore interactive Jupyter notebooks:

jupyter notebook notebooks/

Available notebooks:

  • CA10.ipynb - Main comprehensive notebook
  • 01_Advanced_KG_Embeddings.ipynb - Embedding models
  • 02_Graph_Neural_Networks.ipynb - GNN implementations

Step 6: Use the API (Optional)

Start the API Server

python src/start_api.py --kg-type scientific --port 8000

Access the API

Example API Request

import requests

response = requests.post("http://localhost:8000/query", json={
    "query": "What did Einstein discover?",
    "max_results": 5
})

print(response.json())

Common Tasks

1. Create a Simple Knowledge Graph

from src.core.knowledge_graph import KnowledgeGraph, Entity, Relation

# Create graph
kg = KnowledgeGraph()

# Add entities
kg.add_entity(Entity("Python", "programming_language"))
kg.add_entity(Entity("Guido van Rossum", "person"))

# Add relation
kg.add_relation(Relation("Guido van Rossum", "created", "Python"))

# Visualize
kg.visualize()

2. Train an Embedding Model

from src.models.embeddings import TransE, KGEmbeddingTrainer

# Initialize model
model = TransE(num_entities=100, num_relations=10, embedding_dim=50)

# Train
trainer = KGEmbeddingTrainer(model, kg)
trainer.train(epochs=100)

# Evaluate
results = trainer.evaluate_link_prediction(k=3)
print(f"Hits@3: {results['hits@3']:.3f}")

3. Perform Reasoning

from src.reasoning.langgraph_reasoning import KnowledgeGraphReasoner

# Initialize reasoner
reasoner = KnowledgeGraphReasoner(kg)

# Reason about a query
result = await reasoner.reason("What are Einstein's contributions?")
print(result['final_answer'])

Troubleshooting

Issue: Import Errors

Problem: ModuleNotFoundError: No module named 'src'

Solution: Make sure you're running from the project root:

cd /path/to/CA10_knowledge_graphs
python -c "import sys; sys.path.append('.'); from src.core.knowledge_graph import KnowledgeGraph"

Issue: API Key Errors

Problem: Error: GOOGLE_API_KEY not found

Solution: Set up your .env file with valid API keys or run without advanced features.

Issue: Memory Errors

Problem: RuntimeError: CUDA out of memory

Solution: Reduce batch size or embedding dimensions:

model = TransE(embedding_dim=32)  # Smaller dimension
trainer.train(batch_size=16)       # Smaller batch

Issue: Slow Execution

Problem: Training takes too long

Solution:

  • Reduce number of epochs
  • Use GPU acceleration
  • Enable batch processing

Next Steps

Now that you're set up, explore:

  1. Documentation: Read docs/README.md for comprehensive documentation
  2. Advanced Features: Check docs/README_ADVANCED.md for advanced capabilities
  3. Demos: Explore demos/ for real-world examples
  4. Tests: Run tests/ to understand the codebase

Getting Help

  • Documentation: Check the docs/ folder
  • Examples: Look at demos/ for working examples
  • Issues: Report bugs or ask questions on the repository
  • Community: Join discussions and share your work

Quick Reference

Project Structure

CA10_knowledge_graphs/
โ”œโ”€โ”€ notebooks/     # Jupyter notebooks
โ”œโ”€โ”€ docs/          # Documentation
โ”œโ”€โ”€ scripts/       # Execution scripts
โ”œโ”€โ”€ src/           # Source code
โ”œโ”€โ”€ tests/         # Tests
โ”œโ”€โ”€ config/        # Configuration
โ”œโ”€โ”€ data/          # Results and logs
โ”œโ”€โ”€ models/        # Saved models
โ”œโ”€โ”€ integrations/  # External integrations
โ””โ”€โ”€ demos/         # Demo projects

Key Commands

# Run project
cd scripts && ./run.sh

# Install dependencies
pip install -r config/requirements.txt

# Start API
python src/start_api.py

# Run tests
cd tests && python test_setup.py

# Open notebooks
jupyter notebook notebooks/

Ready to build amazing knowledge graphs! ๐Ÿš€

For more detailed information, see the complete documentation.

Xet Storage Details

Size:
6.71 kB
ยท
Xet hash:
627ea9b3aa1c493429fe719cca410b79882e302eb3c928ba7f41047ecb70bc2e

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.