Buckets:
| # CA10 Knowledge Graphs - Complete Documentation | |
| ## ๐ Overview | |
| This project implements an advanced knowledge graph system with comprehensive reasoning capabilities, embedding models, and modern AI integration. It features LangGraph-based reasoning, RAG (Retrieval-Augmented Generation) systems, and FastAPI-based APIs for interactive knowledge graph exploration. | |
| ## ๐ฏ Key Features | |
| ### Core Knowledge Graph Capabilities | |
| - **Knowledge Graph Construction**: Building graphs from structured and unstructured data | |
| - **Entity Recognition**: Identifying and extracting entities from text | |
| - **Relation Extraction**: Discovering relationships between entities | |
| - **Graph Embeddings**: Learning vector representations of entities and relations | |
| - **Graph Reasoning**: Performing logical inference over knowledge graphs | |
| ### Advanced AI Integration | |
| - **LangGraph Reasoning**: Multi-agent reasoning workflows using LangGraph | |
| - **RAG System**: Retrieval-Augmented Generation with vector search | |
| - **Gemini API Integration**: Advanced reasoning with Google's Gemini | |
| - **Multi-Agent Systems**: Coordinated reasoning across multiple agents | |
| - **Vector Search**: Semantic search using ChromaDB and FAISS | |
| ### Modern Architecture | |
| - **FastAPI Backend**: RESTful API for knowledge graph operations | |
| - **Real-time Visualization**: Interactive graph visualization | |
| - **Scalable Processing**: Efficient processing of large knowledge graphs | |
| - **Comprehensive Evaluation**: Multiple evaluation metrics and benchmarks | |
| ## ๐ Quick Start | |
| ### Prerequisites | |
| - Python 3.8+ | |
| - pip package manager | |
| - Optional: Google Gemini API key (for advanced features) | |
| - Optional: CUDA-capable GPU (for faster training) | |
| ### Installation | |
| ```bash | |
| # Navigate to project directory | |
| cd /path/to/CA10_knowledge_graphs | |
| # Install dependencies | |
| pip install -r config/requirements.txt | |
| # Run the project | |
| cd scripts | |
| ./run.sh | |
| ``` | |
| ### Configuration | |
| Create a `.env` file in the project root with your API keys: | |
| ```bash | |
| # Copy the example configuration | |
| cp config/env_example.txt .env | |
| # Edit .env with your API keys | |
| GOOGLE_API_KEY=your_gemini_api_key_here | |
| OPENAI_API_KEY=your_openai_key_here | |
| ANTHROPIC_API_KEY=your_anthropic_key_here | |
| ``` | |
| ## ๐ Usage Guide | |
| ### Basic Knowledge Graph Operations | |
| ```python | |
| from src.core.knowledge_graph import KnowledgeGraph, Entity, Relation | |
| # Create knowledge graph | |
| kg = KnowledgeGraph() | |
| # Add entities | |
| kg.add_entity(Entity("Einstein", "person")) | |
| kg.add_entity(Entity("Physics", "field")) | |
| # Add relations | |
| kg.add_relation(Relation("Einstein", "studied", "Physics")) | |
| # Query the graph | |
| results = kg.query("What did Einstein study?") | |
| print(f"Results: {results}") | |
| ``` | |
| ### Advanced Scientific Knowledge Graph | |
| ```python | |
| from src.core.scientific_kg import AdvancedScientificKG | |
| # Create scientific knowledge graph | |
| scientific_kg = AdvancedScientificKG() | |
| # Build from scientific data | |
| scientific_kg.build_from_data(scientific_data) | |
| # Analyze influence network | |
| influence_scores = scientific_kg.analyze_influence() | |
| print(f"Influence scores: {influence_scores}") | |
| ``` | |
| ### Embedding Model Training | |
| ```python | |
| from src.models.embeddings import TransE, KGEmbeddingTrainer | |
| # Initialize TransE model | |
| transe = TransE(embedding_dim=100) | |
| # Initialize trainer | |
| trainer = KGEmbeddingTrainer(transe, kg) | |
| # Train the model | |
| trainer.train(num_epochs=100) | |
| # Evaluate | |
| results = trainer.evaluate_link_prediction(k=3) | |
| print(f"Hits@3: {results['hits@k']:.3f}") | |
| ``` | |
| ### LangGraph Reasoning | |
| ```python | |
| from src.reasoning.langgraph_reasoning import KnowledgeGraphReasoner | |
| # Initialize reasoner | |
| reasoner = KnowledgeGraphReasoner(kg) | |
| # Perform reasoning | |
| result = await reasoner.reason("What are the connections between Einstein and quantum mechanics?") | |
| print(f"Reasoning result: {result['final_answer']}") | |
| print(f"Confidence: {result['confidence']:.3f}") | |
| ``` | |
| ### RAG System | |
| ```python | |
| from src.reasoning.rag_system import KnowledgeGraphRAG | |
| # Initialize RAG system | |
| rag = KnowledgeGraphRAG(kg) | |
| # Query with RAG | |
| result = await rag.query("Explain Einstein's contributions to physics") | |
| print(f"RAG answer: {result['answer']}") | |
| print(f"Retrieved documents: {len(result['retrieved_documents'])}") | |
| ``` | |
| ### API Usage | |
| ```python | |
| import requests | |
| # Query via API | |
| response = requests.post("http://localhost:8000/query", json={ | |
| "query": "What did Einstein discover?", | |
| "max_results": 5, | |
| "include_reasoning": True | |
| }) | |
| result = response.json() | |
| print(f"API response: {result['answer']}") | |
| ``` | |
| ## ๐๏ธ Architecture | |
| ### System Components | |
| ``` | |
| โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ | |
| โ CA10 Knowledge Graphs โ | |
| โโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโค | |
| โ Core KG โ Embeddings โ Reasoning โ API โ | |
| โ System โ Models โ Engine โ Server โ | |
| โโโโโโโโโโโโโโโดโโโโโโโโโโโโโโดโโโโโโโโโโโโโโดโโโโโโโโโโโโโโ | |
| ``` | |
| ### Data Flow | |
| ``` | |
| Input Data โ Entity/Relation Extraction โ Knowledge Graph Construction | |
| โ | |
| Embedding Training | |
| โ | |
| Reasoning & Inference | |
| โ | |
| API/Visualization | |
| ``` | |
| ## ๐ Evaluation Metrics | |
| ### Embedding Models | |
| - **Hits@K**: Proportion of correct entities in top-K predictions | |
| - **Mean Rank**: Average rank of correct entities | |
| - **Mean Reciprocal Rank (MRR)**: Average of reciprocal ranks | |
| ### Reasoning Tasks | |
| - **Accuracy**: Percentage of correct inferences | |
| - **Precision/Recall**: Quality of retrieved information | |
| - **F1 Score**: Harmonic mean of precision and recall | |
| ## ๐ง Advanced Configuration | |
| ### Model Parameters | |
| ```python | |
| # Embedding configuration | |
| EMBEDDING_CONFIG = { | |
| 'embedding_dim': 100, | |
| 'learning_rate': 0.001, | |
| 'num_epochs': 100, | |
| 'batch_size': 32, | |
| 'negative_samples': 10 | |
| } | |
| # Reasoning configuration | |
| REASONING_CONFIG = { | |
| 'max_reasoning_steps': 10, | |
| 'confidence_threshold': 0.7, | |
| 'beam_size': 5, | |
| 'temperature': 0.1 | |
| } | |
| ``` | |
| ### Environment Variables | |
| ```bash | |
| # API Configuration | |
| API_HOST=0.0.0.0 | |
| API_PORT=8000 | |
| RAG_API_PORT=8001 | |
| # Database Configuration | |
| NEO4J_URI=bolt://localhost:7687 | |
| NEO4J_USER=neo4j | |
| NEO4J_PASSWORD=your_password | |
| # ChromaDB Configuration | |
| CHROMA_PERSIST_DIRECTORY=./chroma_db | |
| # Logging Configuration | |
| LOG_LEVEL=INFO | |
| LOG_FILE=data/logs/kg_execution.log | |
| ``` | |
| ## ๐งช Testing | |
| ### Running Tests | |
| ```bash | |
| # Navigate to tests directory | |
| cd tests | |
| # Run all tests | |
| python test_setup.py | |
| python test_api_keys.py | |
| # Run with pytest | |
| pytest test_*.py | |
| ``` | |
| ### Test Coverage | |
| - **Unit Tests**: Core functionality testing | |
| - **Integration Tests**: Component interaction testing | |
| - **API Tests**: Endpoint testing | |
| - **Performance Tests**: Scalability and speed testing | |
| ## ๐ Performance Optimization | |
| ### Best Practices | |
| 1. **Batch Processing**: Process multiple entities/relations at once | |
| 2. **Caching**: Cache frequently accessed data | |
| 3. **Parallel Processing**: Use multi-threading for independent operations | |
| 4. **GPU Acceleration**: Enable CUDA for model training | |
| ### Memory Management | |
| ```python | |
| # Use generators for large datasets | |
| def process_large_dataset(data): | |
| for batch in data.batch(size=1000): | |
| yield process_batch(batch) | |
| # Clear cache periodically | |
| kg.clear_cache() | |
| ``` | |
| ## ๐ Troubleshooting | |
| ### Common Issues | |
| **Issue**: API key errors | |
| - **Solution**: Ensure all required API keys are set in `.env` file | |
| **Issue**: Dependency conflicts | |
| - **Solution**: Run `pip install -r config/requirements.txt --upgrade` | |
| **Issue**: Memory errors during training | |
| - **Solution**: Reduce batch size or use smaller embedding dimensions | |
| **Issue**: Slow inference | |
| - **Solution**: Enable GPU acceleration or use cached results | |
| ### Debug Mode | |
| ```bash | |
| # Run with debug logging | |
| export LOG_LEVEL=DEBUG | |
| python src/main_execution.py | |
| ``` | |
| ## ๐ Additional Resources | |
| ### Documentation Files | |
| - **README_RUN.md**: Running instructions | |
| - **README_ADVANCED.md**: Advanced features | |
| - **EXECUTION_GUIDE.md**: Step-by-step guide | |
| ### External Resources | |
| - [Knowledge Graph Papers](https://github.com/topics/knowledge-graph) | |
| - [Graph Neural Networks](https://distill.pub/2021/gnn-intro/) | |
| - [LangGraph Documentation](https://langchain-ai.github.io/langgraph/) | |
| ## ๐ค Contributing | |
| We welcome contributions! Please follow these guidelines: | |
| 1. Fork the repository | |
| 2. Create a feature branch | |
| 3. Make your changes | |
| 4. Add tests for new features | |
| 5. Submit a pull request | |
| ## ๐ License | |
| This project is part of the System2_in_AI CA collection. | |
| --- | |
| **Last Updated**: January 2025 | |
| **Version**: 2.0.0 | |
| **Maintainer**: AI Systems Course Team |
Xet Storage Details
- Size:
- 9.12 kB
- Xet hash:
- ed5a429c32181974f87392ddc31722593203a355209f234c0de616cecbcc4432
ยท
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.