options-recognizer-gliner2-large

LoRA adapter (v4) for fastino/gliner2-large-v1 (486M parameters) that extracts option-contract details from tweets: ticker, strike, option type, expiry, premium and price.

Smaller bases are faster, larger ones are more accurate — see the sibling options-recognizer-* repos for the other sizes.

Code, data pipeline and benchmark: StephanAkkerman/options-recognizer-model.

Intended uses

Retail traders post option trades and unusual-options-flow alerts in a terse, slang-heavy format ($AAPL 350 C 11/06 $1.2M 4.37avg) that regular NER models and regexes handle poorly. This model turns such text into structured fields, which you can use to:

  • build a feed or dashboard of option flow from X/Twitter or Discord alerts,
  • collect option trades into a dataset for research or backtesting,
  • filter or rank alerts by premium, expiry or ticker,
  • pre-label more data for further training.

It is built for English, informal, tweet-length text. It extracts the pieces of a contract as separate spans; it does not pair them up when a text mentions several contracts, and it is not financial advice.

Entities

label what it captures
ticker The stock or ETF symbol an option contract is written on, usually 1-5 letters and often preceded by a dollar sign (e.g., $AAPL, SPY, $T). MUST NOT be a company name, an @handle, or slang (YOLO, NFA, OTM, ITM).
strike The strike price of an option contract, e.g. 350, 162.5, or $1300 in '$1300 calls'. MUST NOT be the stock's current price, a premium or per-contract price, a date, or a percentage.
option_type Whether the contract is a call or a put, in any written form: C, P, call, calls, put, puts, c, p. Only the call/put word itself, not words like buyer or seller.
expiry The expiration date of an option contract: 10/16, 11/06/2026, 6/17/27, November 20, 2026, December, or relative like '2 days'. MUST NOT be the date a report was posted (e.g. the 10/2 in '10/2 Notable Flow').
premium The total dollar size of the trade, e.g. $1.2M, $993K, 23 million. MUST NOT be a strike, the per-contract price, or the stock price.
price The per-contract price or average fill, e.g. .15 in '@ .15' or 4.37 in '4.37avg'. Only the number. MUST NOT be a strike, premium, or the underlying's stock price.

Examples

Actual output of this model (threshold 0.75); empty labels are omitted.

Input:  $AAPL 350 C 11/06/2026 $1.2M 4.37avg
Output: {"ticker": ["$AAPL"], "strike": ["350"], "option_type": ["C"], "expiry": ["11/06/2026"], "premium": ["$1.2M"], "price": ["4.37"]}
Input:  $NVTS - $162K Call buyer
Output: {"ticker": ["$NVTS"], "option_type": ["Call"], "premium": ["$162K"]}
Input:  Bought SPY 600 puts expiring 12/19 @ .85, bearish into CPI
Output: {"ticker": ["SPY"], "strike": ["600"], "option_type": ["puts"], "expiry": ["12/19"], "price": [".85"]}
Input:  Unusual flow: $TSLA $300 calls 01/16/2027 sweep, $2.5M premium, filled at 12.40
Output: {"ticker": ["$TSLA"], "strike": ["$300"], "option_type": ["calls"], "expiry": ["01/16/2027"], "premium": ["$2.5M"], "price": ["12.40"]}
Input:  Just a normal tweet about the market today, no trades.
Output: {}

Results (exact-span match, held-out test set 9bbc0005d8a592af)

label precision recall F1
ticker 97.2% 100.0% 98.6%
strike 98.2% 99.7% 99.0%
option_type 99.6% 100.0% 99.8%
expiry 99.0% 97.9% 98.4%
premium 98.3% 100.0% 99.2%
price 95.1% 98.7% 96.9%
overall 98.0% 99.5% 98.7%

Inference: 95.93 ms/doc (batched, cuda).

Choosing a model

All sizes are trained on the same data and scored on the same test set.

model base params overall F1 ms/doc
GLiNER2.5 Small v1 74M 86.9% 11.95
GLiNER2.5 Base v1 194M 86.8% 10.49
GLiNER2 Base v1 208M 98.5% 12.38
GLiNER2 Large v1 (this repo) 486M 98.7% 95.93

Usage

Needs pip install "gliner2>=2" huggingface_hub.

import json
import re

from gliner2 import AutoExtractor
from huggingface_hub import snapshot_download

# AlnumBoundarySplitter: copy the class from the section below

adapter_dir = snapshot_download("StephanAkkerman/options-recognizer-gliner2-large")
cfg = json.load(open(f"{adapter_dir}/recognizer_config.json"))
model = AutoExtractor.from_pretrained(cfg["base_model"])
model.load_adapter(adapter_dir)
model.set_word_splitter(AlnumBoundarySplitter())  # required, see below

result = model.extract_entities(
    "$AAPL 350 C 11/06/2026 $1.2M 4.37avg",
    cfg["entity_descriptions"],
    threshold=cfg["threshold"],
)
print(result)

Recommended threshold: 0.75. Lower it to find more entities (higher recall), raise it for fewer, more certain ones.

Required word splitter

The adapter was trained with a custom word splitter that breaks between letters and digits, so 1.58avg and 11/20exp tokenize cleanly. Define it before loading the adapter, or accuracy drops sharply:

import re

class AlnumBoundarySplitter:
    """GLiNER2 word splitter that also breaks between letters and digits.

    The stock splitter keeps whole words together, so in ``1.58avg``, ``11/20exp``
    and ``375C`` the gold span ends mid-word and can neither be trained on nor
    predicted. Splitting at letter/digit boundaries makes those spans
    word-aligned. Dates (``10/16/2026``) and numbers with decimals or thousands
    separators (``5.1``, ``18,140``) stay whole so a fragment such as the ``10``
    of a date can never be proposed as a strike. Must be used identically at
    train and inference time.
    """

    _PATTERN = re.compile(
        r"""(?:https?://[^\s]+|www\.[^\s]+)
        |[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}
        |@[a-z0-9_]+
        |[^\W\d_]+(?:[-_][^\W\d_]+)*
        |\d{1,2}/\d{1,2}(?:/\d{2,4})?(?![\d,.]\d)
        |\d+(?:[.,]\d+)*
        |\S""",
        re.VERBOSE | re.IGNORECASE,
    )

    def __call__(self, text, lower=True):
        for m in self._PATTERN.finditer(text):
            token = m.group()
            yield (token.lower() if lower else token), m.start(), m.end()

Training data

StephanAkkerman/options-ner: about 1,300 tweets about stock options, pre-labeled with an LLM and reviewed by hand in Label Studio, plus the held-out test set of 227 tweets used for the scores above (the dataset's test split). LoRA (rank 32) on the attention and dense layers, with early stopping. The full pipeline is in the GitHub repository.

Limitations

  • Scores are for tweets in the style of the training data; other text (news, filings, non-English) is untested.
  • premium (total trade size) and price (per-contract price) can be confused, especially in the smaller models.
  • The GLiNER2.5 variants score noticeably lower than the GLiNER2 ones.

Citation

@misc{akkerman_options_recognizer,
  author = {Stephan Akkerman},
  title  = {options-recognizer-model},
  year   = {2026},
  url    = {https://github.com/StephanAkkerman/options-recognizer-model}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for StephanAkkerman/options-recognizer-gliner2-large

Adapter
(19)
this model

Dataset used to train StephanAkkerman/options-recognizer-gliner2-large