|
Download README.md from Himasree/CodeEmbed-checkpoints: direct link, hf CLI and curl.
- Browser
- Download file 5.38 kB
-
https://huggingface.co/Himasree/CodeEmbed-checkpoints/resolve/main/README.md
- Command line
-
hf download hf://Himasree/CodeEmbed-checkpoints/README.md
-
curl -L -o README.md https://huggingface.co/Himasree/CodeEmbed-checkpoints/resolve/main/README.md
5.38 kB
| library_name: pytorch | |
| tags: | |
| - code-retrieval | |
| - code-search | |
| - code-embedding | |
| - information-retrieval | |
| - transformer | |
| - bi-encoder | |
| - bm25 | |
| - faiss | |
| - codesearchnet | |
| language: | |
| - en | |
| pipeline_tag: feature-extraction | |
| license: other | |
| # CodeEmbed | |
| ## Overview | |
| CodeEmbed | |
| The project explores neural code retrieval using transformer-based bi-encoder architectures, dense semantic retrieval, BM25 lexical retrieval, and hybrid retrieval strategies. | |
| This repository contains trained model checkpoints from the CodeEmbed experiments, including baseline models, dual-encoder experiments, ablation studies, hybrid retrieval experiments, and Phase 7 evaluation models. | |
| The CodeEmbed training experiments, retrieval system implementation, experiment configurations, evaluation workflow, and associated checkpoints were developed as part of this project. | |
| ## Project Goals | |
| CodeEmbed investigates effective retrieval of source-code functions from natural-language queries. | |
| The project explores: | |
| - Transformer-based code representations | |
| - Bi-encoder architectures | |
| - Dense vector retrieval | |
| - BM25 lexical retrieval | |
| - Hybrid dense and lexical retrieval | |
| - Pooling strategies | |
| - Temperature experiments | |
| - Ablation studies | |
| - FAISS-based retrieval | |
| - Ranking-based retrieval evaluation | |
| ## Dataset | |
| The experiments use an AST-cleaned corpus derived from CodeSearchNet-based data. | |
| The underlying dataset remains subject to its original licensing and attribution requirements. | |
| ## Retrieval Approach | |
| CodeEmbed supports hybrid retrieval by combining dense semantic retrieval, BM25 lexical retrieval, and convex interpolation of retrieval scores. | |
| One demonstrated configuration uses: | |
| - Dense weight (alpha): 0.70 | |
| - BM25 contribution: 0.30 | |
| - Corpus size: 19,632 | |
| ## Checkpoints | |
| The repository contains checkpoints from multiple CodeEmbed experiments. | |
| ### Baseline and Dual Encoder Experiments | |
| - basic/ | |
| - dual/ | |
| - dual_bm25_hard/ | |
| - dual_fixed/ | |
| - dual_fixed_bm25/ | |
| ### Ablation Experiments | |
| - ablation/ | |
| - ablation_basic_sweep/ | |
| - ablation_dual_sweep/ | |
| - ablation_dual_fixed_sweep/ | |
| These experiments investigate different pooling strategies and temperature values. | |
| ### Shared Models | |
| - shared/ | |
| - shared_large/ | |
| ### Phase 7 Models | |
| The Phase 7 experiments include CodeBERT, Jina code model, MiniLM, and UniXcoder. | |
| The checkpoints are located under phase7/. | |
| ## Phase 7 Evaluation | |
| A CodeEmbed Phase 7 CodeBERT experiment was evaluated on a test set containing 19,632 query-code pairs. | |
| | Metric | Score | | |
| | --- | ---: | | |
| | Test MRR | 0.7658 | | |
| | Recall@1 | 0.6890 | | |
| | Recall@5 | 0.8640 | | |
| | Recall@10 | 0.9060 | | |
| | NDCG@10 | 0.7976 | | |
| The reported validation MRR for the first epoch was 0.7356. | |
| These results are specific to the project's evaluation setup and dataset. | |
| ## Hybrid Retrieval Results | |
| A demonstrated CodeEmbed hybrid retrieval configuration using convex interpolation achieved: | |
| - Test MRR: 0.6612 | |
| - Improvement over the corresponding dense configuration: 40.7% | |
| - Search latency: approximately 216.81 ms | |
| - Corpus size: 19,632 | |
| - Dense weight: 0.70 | |
| - Observed leakage rate: 1.23% | |
| These values are specific to the project's evaluation and demonstration configuration. | |
| ## Training Checkpoints | |
| The repository preserves both selected model checkpoints and intermediate training checkpoints. | |
| Files named `best_*.pt` generally represent the selected checkpoint for an experiment. | |
| Files named `checkpoint_epoch_*.pt` and `checkpoint_step_*.pt` represent intermediate training checkpoints. | |
| Intermediate checkpoints are retained to preserve the experimental history. | |
| ## Third-Party Models | |
| Some experiments use or evaluate pretrained or externally developed models, including CodeBERT, UniXcoder, MiniLM, and Jina code models. | |
| The inclusion of these checkpoints does not imply ownership of the underlying pretrained models, architectures, tokenizers, or original training data. | |
| Third-party models and components remain subject to their respective licenses and terms. | |
| ## Attribution and Provenance | |
| CodeEmbed and the associated training experiments, retrieval implementation, evaluation workflow, and project-specific checkpoints were developed by Himasree Panku. | |
| This Hugging Face repository preserves the trained checkpoint artifacts and provides a versioned record of the uploaded files. | |
| Third-party models, datasets, libraries, and other external components remain subject to their original licenses and attribution requirements. | |
| ## Reproducibility | |
| For complete reproduction, users should obtain the corresponding CodeEmbed source code, configuration files, preprocessing pipeline, tokenizer and model configuration, and dependency environment. | |
| ## Repository Structure | |
| - ablation/ | |
| - ablation_basic_sweep/ | |
| - ablation_dual_fixed_sweep/ | |
| - ablation_dual_sweep/ | |
| - basic/ | |
| - dual/ | |
| - dual_bm25_hard/ | |
| - dual_fixed/ | |
| - dual_fixed_bm25/ | |
| - phase7/codebert/ | |
| - phase7/jina_v2_code/ | |
| - phase7/minilm_l6/ | |
| - phase7/unixcoder/ | |
| - shared/ | |
| - shared_large/ | |
| ## License | |
| The licensing of individual checkpoints may depend on the underlying pretrained models and third-party components used to produce them. | |
| No license is granted here for third-party pretrained models, datasets, or other components beyond the rights provided by their respective licenses. | |
| Before redistributing or licensing individual checkpoints, users should verify the applicable third-party licenses and permissions. |