Download sdg/preprocessing/dedupe/__init__.py from fzzhang/svd-code: direct link, hf CLI and curl.
- Browser
- Download file 937 Bytes
-
https://huggingface.co/fzzhang/svd-code/resolve/main/sdg/preprocessing/dedupe/__init__.py
- Command line
-
hf download hf://fzzhang/svd-code/sdg/preprocessing/dedupe/__init__.py
-
curl -L -o __init__.py https://huggingface.co/fzzhang/svd-code/resolve/main/sdg/preprocessing/dedupe/__init__.py
937 Bytes
| """Deduplication package: MinHash near-dup + semantic paraphrase dedup. | |
| Standard SFT dedup pipeline order: | |
| exact (caller's responsibility) | |
| | | |
| v | |
| MinHashDeduplicator - lexical near-dups (formatting variants, edits) | |
| | | |
| v | |
| SemanticDeduplicator - paraphrases (different wording, same meaning) | |
| Both classes operate on a list[str] and return list[int] of indices to KEEP, | |
| so callers re-index back into their own record list. An optional | |
| `key_fn(idx) -> sortable` controls which member of a duplicate cluster wins | |
| (smaller key wins). Default: smallest original index ("first seen"). | |
| """ | |
| from .minhash import MinHashDeduplicator | |
| from .semantic import SemanticDeduplicator | |
| from .templates import find_common_prefixes, strip_template, template_hit_stats | |
| __all__ = [ | |
| "MinHashDeduplicator", | |
| "SemanticDeduplicator", | |
| "find_common_prefixes", | |
| "strip_template", | |
| "template_hit_stats", | |
| ] | |