Instructions to use StephanAkkerman/options-recognizer-gliner2-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use StephanAkkerman/options-recognizer-gliner2-large with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("StephanAkkerman/options-recognizer-gliner2-large") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
options-recognizer-gliner2-large
LoRA adapter (v4) for fastino/gliner2-large-v1 (486M parameters) that extracts option-contract details from tweets: ticker, strike, option type, expiry, premium and price.
Smaller bases are faster, larger ones are more accurate — see the sibling options-recognizer-* repos for the other sizes.
Code, data pipeline and benchmark: StephanAkkerman/options-recognizer-model.
Intended uses
Retail traders post option trades and unusual-options-flow alerts in a terse, slang-heavy format ($AAPL 350 C 11/06 $1.2M 4.37avg) that regular NER models and regexes handle poorly. This model turns such text into structured fields, which you can use to:
- build a feed or dashboard of option flow from X/Twitter or Discord alerts,
- collect option trades into a dataset for research or backtesting,
- filter or rank alerts by premium, expiry or ticker,
- pre-label more data for further training.
It is built for English, informal, tweet-length text. It extracts the pieces of a contract as separate spans; it does not pair them up when a text mentions several contracts, and it is not financial advice.
Entities
| label | what it captures |
|---|---|
ticker |
The stock or ETF symbol an option contract is written on, usually 1-5 letters and often preceded by a dollar sign (e.g., $AAPL, SPY, $T). MUST NOT be a company name, an @handle, or slang (YOLO, NFA, OTM, ITM). |
strike |
The strike price of an option contract, e.g. 350, 162.5, or $1300 in '$1300 calls'. MUST NOT be the stock's current price, a premium or per-contract price, a date, or a percentage. |
option_type |
Whether the contract is a call or a put, in any written form: C, P, call, calls, put, puts, c, p. Only the call/put word itself, not words like buyer or seller. |
expiry |
The expiration date of an option contract: 10/16, 11/06/2026, 6/17/27, November 20, 2026, December, or relative like '2 days'. MUST NOT be the date a report was posted (e.g. the 10/2 in '10/2 Notable Flow'). |
premium |
The total dollar size of the trade, e.g. $1.2M, $993K, 23 million. MUST NOT be a strike, the per-contract price, or the stock price. |
price |
The per-contract price or average fill, e.g. .15 in '@ .15' or 4.37 in '4.37avg'. Only the number. MUST NOT be a strike, premium, or the underlying's stock price. |
Examples
Actual output of this model (threshold 0.75); empty labels are omitted.
Input: $AAPL 350 C 11/06/2026 $1.2M 4.37avg
Output: {"ticker": ["$AAPL"], "strike": ["350"], "option_type": ["C"], "expiry": ["11/06/2026"], "premium": ["$1.2M"], "price": ["4.37"]}
Input: $NVTS - $162K Call buyer
Output: {"ticker": ["$NVTS"], "option_type": ["Call"], "premium": ["$162K"]}
Input: Bought SPY 600 puts expiring 12/19 @ .85, bearish into CPI
Output: {"ticker": ["SPY"], "strike": ["600"], "option_type": ["puts"], "expiry": ["12/19"], "price": [".85"]}
Input: Unusual flow: $TSLA $300 calls 01/16/2027 sweep, $2.5M premium, filled at 12.40
Output: {"ticker": ["$TSLA"], "strike": ["$300"], "option_type": ["calls"], "expiry": ["01/16/2027"], "premium": ["$2.5M"], "price": ["12.40"]}
Input: Just a normal tweet about the market today, no trades.
Output: {}
Results (exact-span match, held-out test set 9bbc0005d8a592af)
| label | precision | recall | F1 |
|---|---|---|---|
| ticker | 97.2% | 100.0% | 98.6% |
| strike | 98.2% | 99.7% | 99.0% |
| option_type | 99.6% | 100.0% | 99.8% |
| expiry | 99.0% | 97.9% | 98.4% |
| premium | 98.3% | 100.0% | 99.2% |
| price | 95.1% | 98.7% | 96.9% |
| overall | 98.0% | 99.5% | 98.7% |
Inference: 95.93 ms/doc (batched, cuda).
Choosing a model
All sizes are trained on the same data and scored on the same test set.
| model | base params | overall F1 | ms/doc |
|---|---|---|---|
| GLiNER2.5 Small v1 | 74M | 86.9% | 11.95 |
| GLiNER2.5 Base v1 | 194M | 86.8% | 10.49 |
| GLiNER2 Base v1 | 208M | 98.5% | 12.38 |
| GLiNER2 Large v1 (this repo) | 486M | 98.7% | 95.93 |
Usage
Needs pip install "gliner2>=2" huggingface_hub.
import json
import re
from gliner2 import AutoExtractor
from huggingface_hub import snapshot_download
# AlnumBoundarySplitter: copy the class from the section below
adapter_dir = snapshot_download("StephanAkkerman/options-recognizer-gliner2-large")
cfg = json.load(open(f"{adapter_dir}/recognizer_config.json"))
model = AutoExtractor.from_pretrained(cfg["base_model"])
model.load_adapter(adapter_dir)
model.set_word_splitter(AlnumBoundarySplitter()) # required, see below
result = model.extract_entities(
"$AAPL 350 C 11/06/2026 $1.2M 4.37avg",
cfg["entity_descriptions"],
threshold=cfg["threshold"],
)
print(result)
Recommended threshold: 0.75. Lower it to find more entities (higher recall), raise it for fewer, more certain ones.
Required word splitter
The adapter was trained with a custom word splitter that breaks between letters and digits, so 1.58avg and 11/20exp tokenize cleanly. Define it before loading the adapter, or accuracy drops sharply:
import re
class AlnumBoundarySplitter:
"""GLiNER2 word splitter that also breaks between letters and digits.
The stock splitter keeps whole words together, so in ``1.58avg``, ``11/20exp``
and ``375C`` the gold span ends mid-word and can neither be trained on nor
predicted. Splitting at letter/digit boundaries makes those spans
word-aligned. Dates (``10/16/2026``) and numbers with decimals or thousands
separators (``5.1``, ``18,140``) stay whole so a fragment such as the ``10``
of a date can never be proposed as a strike. Must be used identically at
train and inference time.
"""
_PATTERN = re.compile(
r"""(?:https?://[^\s]+|www\.[^\s]+)
|[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}
|@[a-z0-9_]+
|[^\W\d_]+(?:[-_][^\W\d_]+)*
|\d{1,2}/\d{1,2}(?:/\d{2,4})?(?![\d,.]\d)
|\d+(?:[.,]\d+)*
|\S""",
re.VERBOSE | re.IGNORECASE,
)
def __call__(self, text, lower=True):
for m in self._PATTERN.finditer(text):
token = m.group()
yield (token.lower() if lower else token), m.start(), m.end()
Training data
StephanAkkerman/options-ner: about 1,300 tweets about stock options, pre-labeled with an LLM and reviewed by hand in Label Studio, plus the held-out test set of 227 tweets used for the scores above (the dataset's test split). LoRA (rank 32) on the attention and dense layers, with early stopping. The full pipeline is in the GitHub repository.
Limitations
- Scores are for tweets in the style of the training data; other text (news, filings, non-English) is untested.
premium(total trade size) andprice(per-contract price) can be confused, especially in the smaller models.- The GLiNER2.5 variants score noticeably lower than the GLiNER2 ones.
Citation
@misc{akkerman_options_recognizer,
author = {Stephan Akkerman},
title = {options-recognizer-model},
year = {2026},
url = {https://github.com/StephanAkkerman/options-recognizer-model}
}
Model tree for StephanAkkerman/options-recognizer-gliner2-large
Base model
fastino/gliner2-large-v1