Instructions to use replicate/mrope_get_position_ids with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use replicate/mrope_get_position_ids with Kernels:
# !pip install kernels from kernels import get_kernel # a version (or an explicit revision) is required; see the "Files and versions" tab for the available ones kernel = get_kernel("replicate/mrope_get_position_ids", version=1) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from replicate/mrope_get_position_ids: direct link, hf CLI and curl.
- Browser
- Download file 1.46 kB
-
https://huggingface.co/replicate/mrope_get_position_ids/resolve/main/README.md
- Command line
-
hf download hf://replicate/mrope_get_position_ids/README.md
-
curl -L -o README.md https://huggingface.co/replicate/mrope_get_position_ids/resolve/main/README.md
1.46 kB
metadata
tags:
- kernels
- mrope
Starting from September 13, 2026, we will be removing the "model" type repositories of kernels (e.g., kernels-community/flash-attn3). Make sure you're using a latest version of kernels. If you face any disruption, please report them here: https://github.com/huggingface/kernels/issues/new.
mrope get position ids
This repo is a small rewrite of the get_position_ids function that simply returns the position ids for multimodal input ids.
The goal of this repo is to provide a simplem, close to 1:1 rewrite of the python implementation as a C++ and small CUDA kernel implementation.
nix develop -L
and then running the following command:
pytest test/test.py -s
# platform linux -- Python 3.12.8, pytest-8.3.3, pluggy-1.5.0
# rootdir: /root/mrope_get_position_ids
# collected 3 items
#
# test/test.py
# Vision config one_segment - Reference time: 131.56 ms
# Vision config one_segment - Extension time: 6.76 ms
#
# .
# Vision config two_segments - Reference time: 1.72 ms
# Vision config two_segments - Extension time: 0.56 ms
#
# .
# Vision config three_segments - Reference time: 2.02 ms
# Vision config three_segments - Extension time: 0.49 ms
#
# .
Notes
- this example isn't that expensive overally, but the kernel avoids the need to copy the data back and forth between the CPU and GPU and allocates all of the values in the tensor in parallel.