new

Get trending papers in your email inbox!

Subscribe

Daily Papers

byAK and the research community

Oct 6

rsx: A high-performance streaming toolkit for RAD-seq sex determination

Background Restriction site-associated DNA sequencing (RAD-seq) is widely used to discover sex-linked markers in non-model organisms, and RADSex provides the reference workflow for building marker-by-individual depth tables and testing sex-biased marker distributions. Its table-building commands grow memory-hungry as panels reach millions of RAD tags, it reports frequentist calls with no posterior evidence, and it offers no Python or C interface. Results rsx is a Rust implementation of the complete RADSex command set that preserves marker-table semantics and command-line compatibility. It combines 2-bit DNA keys, parallel ingestion, memory-mapped tables, external sorting, bitset group counts and a streamed Gram matrix so that writable allocations stay bounded by the number of individuals or by an explicit buffer, with false-discovery-rate ranking the one deliberate exception. Conjugate Beta-Binomial Bayes factors and directional posteriors grade each marker as a strict call, a posterior-supported hypothesis or a Bayes-factor-only row, and an optional CUDA backend batches the per-marker arithmetic on the GPU. On four published RAD-seq panels comprising 41.9 billion sequenced bases, rsx reproduced the RADSex v1.2.0 calls, recovered every Bonferroni-significant positive-control marker, and was 8.38-fold faster in geometric mean across 56 paired timings; the CUDA backend adds up to 29.86-fold on the p-value batch. Python and C bindings drive the same core from notebooks and pipelines. Conclusions rsx is an allocation-bounded, statistically extended replacement for RADSex that stays backward-compatible and reports its evidence in explicit grades. It is released under the GPL-3.0-or-later licence, with a reproducibility archive covering every reported number.

  • 2 authors
·
Aug 2 1

AutoCoreset: An Automatic Practical Coreset Construction Framework

A coreset is a tiny weighted subset of an input set, that closely resembles the loss function, with respect to a certain set of queries. Coresets became prevalent in machine learning as they have shown to be advantageous for many applications. While coreset research is an active research area, unfortunately, coresets are constructed in a problem-dependent manner, where for each problem, a new coreset construction algorithm is usually suggested, a process that may take time or may be hard for new researchers in the field. Even the generic frameworks require additional (problem-dependent) computations or proofs to be done by the user. Besides, many problems do not have (provable) small coresets, limiting their applicability. To this end, we suggest an automatic practical framework for constructing coresets, which requires (only) the input data and the desired cost function from the user, without the need for any other task-related computation to be done by the user. To do so, we reduce the problem of approximating a loss function to an instance of vector summation approximation, where the vectors we aim to sum are loss vectors of a specific subset of the queries, such that we aim to approximate the image of the function on this subset. We show that while this set is limited, the coreset is quite general. An extensive experimental study on various machine learning applications is also conducted. Finally, we provide a ``plug and play" style implementation, proposing a user-friendly system that can be easily used to apply coresets for many problems. Full open source code can be found at https://github.com/alaamaalouf/AutoCoreset{https://github.com/alaamaalouf/AutoCoreset}. We believe that these contributions enable future research and easier use and applications of coresets.

  • 4 authors
·
May 19, 2023

Coverage-centric Coreset Selection for High Pruning Rates

One-shot coreset selection aims to select a representative subset of the training data, given a pruning rate, that can later be used to train future models while retaining high accuracy. State-of-the-art coreset selection methods pick the highest importance examples based on an importance metric and are found to perform well at low pruning rates. However, at high pruning rates, they suffer from a catastrophic accuracy drop, performing worse than even random sampling. This paper explores the reasons behind this accuracy drop both theoretically and empirically. We first propose a novel metric to measure the coverage of a dataset on a specific distribution by extending the classical geometric set cover problem to a distribution cover problem. This metric helps explain why coresets selected by SOTA methods at high pruning rates perform poorly compared to random sampling because of worse data coverage. We then propose a novel one-shot coreset selection method, Coverage-centric Coreset Selection (CCS), that jointly considers overall data coverage upon a distribution as well as the importance of each example. We evaluate CCS on five datasets and show that, at high pruning rates (e.g., 90%), it achieves significantly better accuracy than previous SOTA methods (e.g., at least 19.56% higher on CIFAR10) as well as random selection (e.g., 7.04% higher on CIFAR10) and comparable accuracy at low pruning rates. We make our code publicly available at https://github.com/haizhongzheng/Coverage-centric-coreset-selection.

  • 4 authors
·
Oct 27, 2022