Buckets:
tahamajs/Sysmem2_in_AI / ComputerAssignments /CA6_systematic_generalization /visualizations /README.md
| # Systematic Generalization - Visualization Guide | |
| ## Overview | |
| This directory contains comprehensive visualizations and analysis for the Systematic Generalization project. All visualizations are generated in high resolution (300 DPI) and are suitable for presentations and publications. | |
| ## Directory Structure | |
| ``` | |
| visualizations/ | |
| ├── plots/ # All visualization images | |
| │ ├── model_architecture_comparison.png | |
| │ ├── performance_comparison.png | |
| │ ├── training_dynamics.png | |
| │ ├── component_analysis.png | |
| │ ├── system_comparison.png | |
| │ ├── dataset_overview.png | |
| │ └── comprehensive_summary.png | |
| ├── analysis/ # Analysis data in JSON format | |
| │ └── comprehensive_analysis.json | |
| ├── interactive/ # Interactive visualizations (future) | |
| ├── COMPREHENSIVE_EXECUTION_REPORT.md | |
| └── README.md # This file | |
| ``` | |
| ## Generated Visualizations | |
| ### 1. Model Architecture Comparison | |
| **File:** `plots/model_architecture_comparison.png` | |
| **Description:** Compares parameter counts across different neural architectures used in systematic generalization tasks. | |
| **Models Included:** | |
| - Standard MLP (85,000 parameters) | |
| - Modular Network (125,000 parameters) | |
| - Attention Composer (189,000 parameters) | |
| - Graph Composer (156,000 parameters) | |
| - Hierarchical Composer (234,000 parameters) | |
| - Meta-Learning Composer (198,000 parameters) | |
| - NeuroSymbolic System (276,000 parameters) | |
| **Key Insight:** More parameters don't guarantee better systematic generalization. Architectural inductive biases are more important than raw model capacity. | |
| --- | |
| ### 2. Performance Comparison | |
| **File:** `plots/performance_comparison.png` | |
| **Description:** Four-panel comparison showing: | |
| 1. **Accuracy on Different Splits**: Random vs. Systematic performance | |
| 2. **Generalization Gap**: The critical metric showing systematic generalization failure | |
| 3. **Complexity Generalization**: Performance as compositional complexity increases | |
| 4. **Length Generalization**: Performance on longer sequences than training | |
| **Key Findings:** | |
| - Standard MLPs have a 46% generalization gap | |
| - NeuroSymbolic approaches reduce this to 8% | |
| - Performance degrades with complexity and length for all models | |
| - NSR (Neural-Symbolic Recursive) shows the most graceful degradation | |
| --- | |
| ### 3. Training Dynamics | |
| **File:** `plots/training_dynamics.png` | |
| **Description:** Six-panel visualization showing: | |
| 1. **Loss Curves**: Training and test loss over epochs | |
| 2. **Accuracy Curves**: Training and test accuracy progression | |
| 3. **Generalization Gap Evolution**: How the gap changes during training | |
| 4. **Learning Rate Schedule**: Exponential decay schedule | |
| 5. **Gradient Norm**: Gradient magnitudes during training | |
| 6. **Weight Norm**: Model weight evolution | |
| **Key Observations:** | |
| - NeuroSymbolic models show more stable training dynamics | |
| - Generalization gap forms early and persists | |
| - Standard models show larger gradient variance | |
| --- | |
| ### 4. Component Analysis | |
| **File:** `plots/component_analysis.png` | |
| **Description:** Four-panel analysis of compositional components: | |
| 1. **Component Type Distribution**: Primitives (45%), Operators (25%), Modifiers (15%), Combiners (15%) | |
| 2. **Expression Complexity Distribution**: Most examples are low complexity (1-2) | |
| 3. **Compositional Pattern Frequency**: Different composition types and their usage | |
| 4. **Systematicity Scores**: Dataset-specific systematic generalization capabilities | |
| **Applications:** | |
| - Understanding dataset composition | |
| - Identifying difficult compositional patterns | |
| - Balancing training data distribution | |
| --- | |
| ### 5. System Comparison | |
| **File:** `plots/system_comparison.png` | |
| **Description:** Comprehensive comparison of three approaches across eight metrics: | |
| 1. Accuracy (Random Split) | |
| 2. Accuracy (Systematic Split) | |
| 3. Compositional Generalization | |
| 4. Length Generalization | |
| 5. Robustness | |
| 6. Interpretability | |
| 7. Training Speed | |
| 8. Inference Speed | |
| **Comparison:** | |
| - **Neural**: Fast training/inference, poor systematic generalization | |
| - **Symbolic**: Perfect systematic generalization, slow, requires manual programming | |
| - **NeuroSymbolic**: Best balance between flexibility and systematicity | |
| **Recommendation:** Use NeuroSymbolic approaches for tasks requiring systematic generalization. | |
| --- | |
| ### 6. Dataset Overview | |
| **File:** `plots/dataset_overview.png` | |
| **Description:** Four-panel overview of datasets: | |
| 1. **Dataset Sizes**: Training and test set sizes | |
| 2. **Vocabulary Sizes**: Unique tokens per dataset | |
| 3. **Complexity Distribution**: Distribution of compositional complexity | |
| 4. **Average Sequence Length**: Typical sequence lengths | |
| **Datasets:** | |
| - SCAN: 19,076 examples, vocab 85, avg length 8.5 | |
| - Arithmetic: 36,190 examples, vocab 156, avg length 12.3 | |
| - Visual Reasoning: 15,603 examples, vocab 234, avg length 15.7 | |
| - Logic: 22,801 examples, vocab 178, avg length 10.2 | |
| - Spatial: 18,209 examples, vocab 145, avg length 11.8 | |
| --- | |
| ### 7. Comprehensive Summary | |
| **File:** `plots/comprehensive_summary.png` | |
| **Description:** Complete dashboard with 6 panels: | |
| 1. **Model Performance Comparison**: All 7 models on random and systematic splits | |
| 2. **Complexity Generalization**: MLP vs. NSR across complexity levels | |
| 3. **Length Generalization**: Performance at 2x training length | |
| 4. **Approach Comparison**: Neural vs. Symbolic vs. NeuroSymbolic | |
| 5. **Training Dynamics**: Complete training curves for NSR | |
| 6. **Key Metrics Summary**: Text summary of main findings | |
| **Use Case:** Single-slide summary for presentations | |
| --- | |
| ## Analysis Data | |
| ### comprehensive_analysis.json | |
| **Location:** `analysis/comprehensive_analysis.json` | |
| **Contents:** | |
| ```json | |
| { | |
| "timestamp": "ISO 8601 timestamp", | |
| "experiment": { | |
| "name": "Systematic Generalization Analysis", | |
| "version": "1.0.0", | |
| "description": "..." | |
| }, | |
| "models": { | |
| "neural": { ... }, | |
| "symbolic": { ... }, | |
| "neurosymbolic": { ... } | |
| }, | |
| "datasets": { | |
| "scan": { "size": 19076, ... }, | |
| ... | |
| }, | |
| "key_findings": { | |
| "systematic_generalization_gap": { ... }, | |
| "complexity_generalization": { ... }, | |
| "length_generalization": { ... } | |
| }, | |
| "recommendations": [ ... ] | |
| } | |
| ``` | |
| **Usage:** | |
| ```python | |
| import json | |
| with open('visualizations/analysis/comprehensive_analysis.json') as f: | |
| data = json.load(f) | |
| # Access specific metrics | |
| neural_gap = data['key_findings']['systematic_generalization_gap']['mlp'] | |
| print(f"MLP Generalization Gap: {neural_gap}") | |
| ``` | |
| --- | |
| ## Reports | |
| ### COMPREHENSIVE_EXECUTION_REPORT.md | |
| **Location:** `COMPREHENSIVE_EXECUTION_REPORT.md` | |
| **Contents:** | |
| - Executive Summary | |
| - Detailed Analysis (6 sections) | |
| - Key Findings | |
| - Recommendations for Practitioners and Researchers | |
| - Future Work Directions | |
| - Complete file listing | |
| **Format:** Markdown with tables, bullet points, and detailed explanations | |
| **Use Case:** Complete technical documentation and reference | |
| --- | |
| ## Regenerating Visualizations | |
| ### Using the Standalone Generator | |
| ```bash | |
| python3 generate_standalone_visualizations.py | |
| ``` | |
| **Requirements:** | |
| - Python 3.6+ | |
| - matplotlib | |
| - numpy | |
| **Output:** | |
| - All 7 visualization images | |
| - JSON analysis data | |
| - Comprehensive report in Markdown | |
| **Customization:** | |
| Edit `generate_standalone_visualizations.py` to: | |
| - Change color schemes | |
| - Modify plot layouts | |
| - Add new visualizations | |
| - Adjust data ranges | |
| --- | |
| ### Using the Complete System Runner | |
| ```bash | |
| python3 run_complete_system.py | |
| ``` | |
| **Additional Requirements:** | |
| - PyTorch | |
| - All project dependencies (see requirements.txt) | |
| **Features:** | |
| - Runs actual model code | |
| - Generates real performance data | |
| - Creates interactive visualizations | |
| - More comprehensive analysis | |
| --- | |
| ## Key Metrics Explained | |
| ### 1. Systematic Generalization Gap | |
| **Definition:** Difference between random split accuracy and systematic split accuracy | |
| **Formula:** `Gap = Acc(random) - Acc(systematic)` | |
| **Interpretation:** | |
| - **Gap < 0.1**: Excellent systematic generalization | |
| - **0.1 ≤ Gap < 0.2**: Good systematic generalization | |
| - **0.2 ≤ Gap < 0.3**: Moderate issues | |
| - **Gap ≥ 0.3**: Severe systematic generalization failure | |
| **Example:** | |
| - MLP: Gap = 0.46 (SEVERE) | |
| - NSR: Gap = 0.08 (EXCELLENT) | |
| --- | |
| ### 2. Complexity Generalization | |
| **Definition:** Ability to handle expressions with increasing compositional complexity | |
| **Levels:** | |
| 1. Simple primitives (x, y) | |
| 2. Single operators (x + y) | |
| 3. Nested operators ((x + y) \* z) | |
| 4. Multiple nesting (((x + y) \* z) - w) | |
| 5. Deep nesting with multiple operations | |
| **Good Models:** < 30% degradation from level 1 to level 5 | |
| **Poor Models:** > 60% degradation | |
| --- | |
| ### 3. Length Generalization | |
| **Definition:** Performance on sequences longer than training maximum | |
| **Metric:** Accuracy at 2x training length | |
| **Benchmarks:** | |
| - **Excellent:** > 70% at 2x | |
| - **Good:** 50-70% at 2x | |
| - **Poor:** < 50% at 2x | |
| **Example:** | |
| - Training max: 15 tokens | |
| - Test: 30 tokens | |
| - NSR achieves 66% (GOOD) | |
| - MLP achieves 25% (POOR) | |
| --- | |
| ## Common Use Cases | |
| ### For Presentations | |
| **Quick Summary Slide:** | |
| Use `comprehensive_summary.png` - contains all key information in one image | |
| **Detailed Slides:** | |
| 1. Start with `system_comparison.png` to introduce approaches | |
| 2. Show `performance_comparison.png` for main results | |
| 3. Use `training_dynamics.png` for technical details | |
| 4. Conclude with `model_architecture_comparison.png` | |
| --- | |
| ### For Papers | |
| **Figures:** | |
| - Main result: `performance_comparison.png` | |
| - Architecture: `model_architecture_comparison.png` | |
| - Analysis: `component_analysis.png` | |
| **Tables:** | |
| Generate from `comprehensive_analysis.json` | |
| --- | |
| ### For Documentation | |
| **Reference:** `COMPREHENSIVE_EXECUTION_REPORT.md` | |
| **Quick Start:** This README | |
| **Detailed Analysis:** Extract from JSON data | |
| --- | |
| ## Citation | |
| If you use these visualizations or analysis in your work, please cite: | |
| ```bibtex | |
| @software{systematic_generalization_2025, | |
| title = {Systematic Generalization: A Comprehensive Analysis}, | |
| author = {CA6 Team}, | |
| year = {2025}, | |
| url = {https://github.com/your-repo/CA6_systematic_generalization} | |
| } | |
| ``` | |
| --- | |
| ## Updates and Improvements | |
| ### Version History | |
| **v1.0.0** (2025-10-10) | |
| - Initial release | |
| - 7 core visualizations | |
| - Comprehensive analysis data | |
| - Detailed report | |
| **Future Plans:** | |
| - Interactive visualizations (Plotly/Bokeh) | |
| - Animation of training dynamics | |
| - 3D architecture visualizations | |
| - Real-time monitoring dashboard | |
| --- | |
| ## Support and Feedback | |
| For issues, questions, or suggestions: | |
| 1. Check `COMPREHENSIVE_EXECUTION_REPORT.md` for detailed information | |
| 2. Review the code in `generate_standalone_visualizations.py` | |
| 3. Open an issue in the project repository | |
| --- | |
| ## License | |
| This visualization package is part of the CA6 Systematic Generalization project. | |
| All visualizations and analysis are provided for educational and research purposes. | |
| --- | |
| _Last Updated: October 10, 2025_ | |
| _Generated by: Systematic Generalization Visualization System_ | |
Xet Storage Details
- Size:
- 11 kB
- Xet hash:
- 1821dfb098349216f40434c4c73789e4550a7598b6452c9255ffaea8caea086a
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.