TommyPanLab commited on
Commit
ec6075a
·
verified ·
1 Parent(s): b0b98a3

Align model card with AAP-SQL GitHub workflow

Browse files
Files changed (1) hide show
  1. README.md +18 -16
README.md CHANGED
@@ -11,39 +11,41 @@ tags:
11
  - aap-sql
12
  ---
13
 
14
- # AAP-SQL E2 column retriever
15
 
16
- AAP-SQL E2 是完整 AAP-SQL 設定中的雙編碼欄位檢索器。它先從資料庫 schema 中取回 50 個候選欄位,再交給 R1 cross-encoder 重排。
17
 
18
- AAP-SQL E2 is the bi-encoder column retriever used by the full AAP-SQL configuration. It retrieves 50 schema-column candidates before the R1 cross-encoder reranks them.
19
 
20
  ## Model details
21
 
22
- - Base model: `sentence-transformers/all-mpnet-base-v2`
23
- - Training objective: `MultipleNegativesRankingLoss`
24
- - Training seed: `42`
25
  - Training data: column-retrieval examples derived from the BIRD training split and schema descriptions
26
- - Expected library: `sentence-transformers>=5.1.2`
 
 
 
 
27
 
28
  ## Use with AAP-SQL
29
 
30
  Download this repository into the path expected by the final runner:
31
 
32
- ```powershell
33
- hf download TommyPanLab/AAP-SQL-E2 --local-dir models/column_retriever_v3_paper_earlystop
34
- ```
35
 
36
  Direct loading:
37
 
38
- ```python
39
  from sentence_transformers import SentenceTransformer
40
 
41
- model = SentenceTransformer("TommyPanLab/AAP-SQL-E2")
42
  embeddings = model.encode(["table.column: column description"])
43
- ```
44
-
45
- The complete pipeline, required BIRD directory layout, and Gemini 3.1 result are documented in the [AAP-SQL publication branch](https://github.com/Tommyweige/AAP-SQL/tree/codex/final-aap-sql-experiment/AAP-SQL-Original).
46
 
47
  ## Data and license notice
48
 
49
- The training examples were derived from the BIRD benchmark. Review the [BIRD project terms](https://bird-bench.github.io/) before using the model. No additional license has been declared for these fine-tuned weights; the upstream model and dataset terms still apply.
 
11
  - aap-sql
12
  ---
13
 
14
+ # AAP-SQL column retriever
15
 
16
+ AAP-SQL 欄位檢索器是完整 AAP-SQL 設定中的雙編碼模型。它從完整資料庫結構與欄位描述中召回最多 50 個候選欄位,供後續重排序器選出核心欄位。
17
 
18
+ AAP-SQL column retriever is the bi-encoder used in the full AAP-SQL workflow. It performs the first stage of schema retrieval and returns up to 50 candidate columns for candidate reranking.
19
 
20
  ## Model details
21
 
22
+ - Base model: sentence-transformers/all-mpnet-base-v2
23
+ - Training objective: MultipleNegativesRankingLoss
24
+ - Training seed: 42
25
  - Training data: column-retrieval examples derived from the BIRD training split and schema descriptions
26
+ - Expected library: sentence-transformers>=5.1.2
27
+
28
+ ## AAP-SQL publication branch
29
+
30
+ The complete AAP-SQL workflow, research method terminology, BIRD directory layout, and reproduction instructions are maintained in the [GitHub publication branch](https://github.com/Tommyweige/AAP-SQL/tree/codex/final-aap-sql-experiment/AAP-SQL-Original).
31
 
32
  ## Use with AAP-SQL
33
 
34
  Download this repository into the path expected by the final runner:
35
 
36
+ ~~~powershell
37
+ hf download TommyPanLab/AAP-SQL-Column-Retriever --local-dir models/column_retriever_v3_paper_earlystop
38
+ ~~~
39
 
40
  Direct loading:
41
 
42
+ ~~~python
43
  from sentence_transformers import SentenceTransformer
44
 
45
+ model = SentenceTransformer("TommyPanLab/AAP-SQL-Column-Retriever")
46
  embeddings = model.encode(["table.column: column description"])
47
+ ~~~
 
 
48
 
49
  ## Data and license notice
50
 
51
+ The training examples were derived from the BIRD benchmark. Review the [BIRD project terms](https://bird-bench.github.io/) before using the model. No additional license has been declared for these fine-tuned weights; the upstream model and dataset terms still apply.