File size: 5,384 Bytes
bb45bb6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cc654c1
 
 
 
b6014d1
cc654c1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
---
library_name: pytorch
tags:
  - code-retrieval
  - code-search
  - code-embedding
  - information-retrieval
  - transformer
  - bi-encoder
  - bm25
  - faiss
  - codesearchnet
language:
  - en
pipeline_tag: feature-extraction
license: other
---

# CodeEmbed

## Overview

CodeEmbed

The project explores neural code retrieval using transformer-based bi-encoder architectures, dense semantic retrieval, BM25 lexical retrieval, and hybrid retrieval strategies.

This repository contains trained model checkpoints from the CodeEmbed experiments, including baseline models, dual-encoder experiments, ablation studies, hybrid retrieval experiments, and Phase 7 evaluation models.


The CodeEmbed training experiments, retrieval system implementation, experiment configurations, evaluation workflow, and associated checkpoints were developed as part of this project.

## Project Goals

CodeEmbed investigates effective retrieval of source-code functions from natural-language queries.

The project explores:

- Transformer-based code representations
- Bi-encoder architectures
- Dense vector retrieval
- BM25 lexical retrieval
- Hybrid dense and lexical retrieval
- Pooling strategies
- Temperature experiments
- Ablation studies
- FAISS-based retrieval
- Ranking-based retrieval evaluation

## Dataset

The experiments use an AST-cleaned corpus derived from CodeSearchNet-based data.

The underlying dataset remains subject to its original licensing and attribution requirements.

## Retrieval Approach

CodeEmbed supports hybrid retrieval by combining dense semantic retrieval, BM25 lexical retrieval, and convex interpolation of retrieval scores.

One demonstrated configuration uses:

- Dense weight (alpha): 0.70
- BM25 contribution: 0.30
- Corpus size: 19,632

## Checkpoints

The repository contains checkpoints from multiple CodeEmbed experiments.

### Baseline and Dual Encoder Experiments

- basic/
- dual/
- dual_bm25_hard/
- dual_fixed/
- dual_fixed_bm25/

### Ablation Experiments

- ablation/
- ablation_basic_sweep/
- ablation_dual_sweep/
- ablation_dual_fixed_sweep/

These experiments investigate different pooling strategies and temperature values.

### Shared Models

- shared/
- shared_large/

### Phase 7 Models

The Phase 7 experiments include CodeBERT, Jina code model, MiniLM, and UniXcoder.

The checkpoints are located under phase7/.

## Phase 7 Evaluation

A CodeEmbed Phase 7 CodeBERT experiment was evaluated on a test set containing 19,632 query-code pairs.

| Metric | Score |
| --- | ---: |
| Test MRR | 0.7658 |
| Recall@1 | 0.6890 |
| Recall@5 | 0.8640 |
| Recall@10 | 0.9060 |
| NDCG@10 | 0.7976 |

The reported validation MRR for the first epoch was 0.7356.

These results are specific to the project's evaluation setup and dataset.

## Hybrid Retrieval Results

A demonstrated CodeEmbed hybrid retrieval configuration using convex interpolation achieved:

- Test MRR: 0.6612
- Improvement over the corresponding dense configuration: 40.7%
- Search latency: approximately 216.81 ms
- Corpus size: 19,632
- Dense weight: 0.70
- Observed leakage rate: 1.23%

These values are specific to the project's evaluation and demonstration configuration.

## Training Checkpoints

The repository preserves both selected model checkpoints and intermediate training checkpoints.

Files named `best_*.pt` generally represent the selected checkpoint for an experiment.

Files named `checkpoint_epoch_*.pt` and `checkpoint_step_*.pt` represent intermediate training checkpoints.

Intermediate checkpoints are retained to preserve the experimental history.

## Third-Party Models

Some experiments use or evaluate pretrained or externally developed models, including CodeBERT, UniXcoder, MiniLM, and Jina code models.

The inclusion of these checkpoints does not imply ownership of the underlying pretrained models, architectures, tokenizers, or original training data.

Third-party models and components remain subject to their respective licenses and terms.

## Attribution and Provenance

CodeEmbed and the associated training experiments, retrieval implementation, evaluation workflow, and project-specific checkpoints were developed by Himasree Panku.

This Hugging Face repository preserves the trained checkpoint artifacts and provides a versioned record of the uploaded files.

Third-party models, datasets, libraries, and other external components remain subject to their original licenses and attribution requirements.

## Reproducibility

For complete reproduction, users should obtain the corresponding CodeEmbed source code, configuration files, preprocessing pipeline, tokenizer and model configuration, and dependency environment.

## Repository Structure

- ablation/
- ablation_basic_sweep/
- ablation_dual_fixed_sweep/
- ablation_dual_sweep/
- basic/
- dual/
- dual_bm25_hard/
- dual_fixed/
- dual_fixed_bm25/
- phase7/codebert/
- phase7/jina_v2_code/
- phase7/minilm_l6/
- phase7/unixcoder/
- shared/
- shared_large/

## License

The licensing of individual checkpoints may depend on the underlying pretrained models and third-party components used to produce them.

No license is granted here for third-party pretrained models, datasets, or other components beyond the rights provided by their respective licenses.

Before redistributing or licensing individual checkpoints, users should verify the applicable third-party licenses and permissions.