UniHive RecSys: hybrid feed-ranking engine (v2.0)

A small, fast post recommendation engine for a student social app. Give it a user (interests + past interactions) and a list of candidate posts; it returns the posts in feed order, with a score and an explanation for each.

βœ… This repository is engine v2.0: our final, intended design. Use hybrid_engine/.

  • Run it: hybrid_engine/ (code) + hybrid_engine/artifacts/ (the weights it loads). Nothing else is needed.
  • Train it further: pytorch_checkpoints/engine_v2.0/ (the same v2.0 models as PyTorch files).
  • Ignore archive/ unless you are studying the project's history. Those are superseded experiments, built on different inputs, and they will not work with hybrid_engine/.

Correction (2026-10-07). Version 2.0 removes the community name and flair as direct inputs and from the embedding text. It still receives each post's domain field, which for Reddit text posts is self.<subreddit>: 24 of the 30 domain tokens known to both models are subreddit names. Community information therefore remains indirectly available to the 2.0 models for text posts. A follow-up release removes domain entirely. Until then, treat v2.0's community independence as partial.


Contents

  1. Which files you need
  2. Install and run in 2 minutes
  3. Use it in your code
  4. Content vectors (recommended)
  5. Inputs and outputs
  6. Settings
  7. Plug it into a backend
  8. Continue training
  9. How it works
  10. Scores
  11. Model size
  12. Where the engine's parts come from
  13. Training data and limitations
  14. Archive (older versions)

1. Which files you need

You want to... Files
Rank posts (most users) hybrid_engine/ (code + artifacts/cold_stage2.onnx, artifacts/warm_deepfm.onnx, artifacts/meta.json)
Turn post text into content vectors the public encoder Qwen/Qwen3-Embedding-0.6B (section 4)
Continue training the models pytorch_checkpoints/engine_v2.0/ (section 8)
Read the full project report REPORT.md

The engine reads exactly the files in hybrid_engine/artifacts/. Don't replace them with files from archive/: the feature lists, vocabularies and standardisation differ.


2. Install and run in 2 minutes

pip install huggingface_hub numpy onnxruntime
hf download Tirthanker/UniHive_RecSys --include "hybrid_engine/*" --local-dir UniHive_RecSys
cd UniHive_RecSys
python -m hybrid_engine.example                 # demo: a brand-new user vs the same user after 30 engaged posts
python -m hybrid_engine.tests.test_engine       # smoke tests; should end with "community / flair guard passed"

Python 3.10+. Runs on CPU; no GPU needed.


3. Use it in your code

import time
from hybrid_engine import HybridRecommender

engine = HybridRecommender()          # load once at startup; reuse for every request

user = {
    "interests": ["machine-learning", "computer-vision"],    # tag slugs the user picked (list: meta.json -> taxonomy.child)
    "history": [                                              # what the user saw before; may be empty for a new user
        {"post_id": "p1", "ts_ms": 1791141039157, "visible_ms": 5200, "focal_ms": 4800, "hover_ms": 300,
         "clicked": True, "engaged": True, "community": "c-ml", "author": "a-17", "embedding": None},
    ],
}
candidates = [
    {"post_id": "p9", "post_type": "text", "score": 42, "comments": 7, "age_hours": 3.5,
     "title": "Fine-tuning ViTs on a laptop GPU", "body": "Notes from a weekend project ...",
     "tags": {"computer-vision": 0.9, "deep-learning": 0.7}, "embedding": None,
     "author": "a-3", "impressions": 0, "community": "c-cv"},
    {"post_id": "p10", "post_type": "image", "score": 300, "comments": 40, "age_hours": 50,
     "title": "Campus fest photos", "body": "", "tags": {"photography": 0.8}, "embedding": None,
     "author": "a-8", "impressions": 120, "community": "c-fest"},
]

feed = engine.recommend(user, candidates, now_ms=time.time() * 1000, k=10)
for d in feed:
    print(d["position"], d["post_id"], round(d["score"], 3), "topic" if d["topic_match"] else "", "warm weight", d["warm_weight"])
  • recommend(user, candidates, now_ms, k): the top k in final feed order.
  • rank(...): every candidate in feed order.
  • score(...): raw scores without the topic-first reordering.

now_ms is the request time in milliseconds. History items at or after it are ignored automatically.


4. Content vectors (recommended)

The models read a 1024-d content vector per post. Build it exactly like this (the same recipe the models were trained on):

from sentence_transformers import SentenceTransformer        # pip install sentence-transformers
from hybrid_engine.backend_adapter import qwen3_text

encoder = SentenceTransformer("Qwen/Qwen3-Embedding-0.6B"); encoder.max_seq_length = 512   # load once
text = qwen3_text(title, body, image_caption)                 # "title. body" (+ " Image: caption"); no community, no flair
post["embedding"] = encoder.encode(text, normalize_embeddings=True)
  • Compute it once per post when the post is created, and store it.
  • Without vectors ("embedding": None) the engine still ranks, but loses the content signal: about 0.02 lower NDCG@10 on average (section 10).
  • A vector of the wrong size (e.g. a 512-d CLIP vector) is ignored on purpose: it lives in a different space.

5. Inputs and outputs

Candidate post

Field Type Meaning
post_id str your id
post_type str text, image, video, link, gallery, ...
score, comments int likes / upvotes and comment count
age_hours float hours since the post was created
title, body, image_caption str used for length features (and for the content vector, section 4)
tags dict or list {slug: confidence} from the 114-tag taxonomy, or a list of slugs (confidence 1.0)
embedding 1024 floats or None section 4
author, community str used only for the user's own history with that author / community, never as model inputs
impressions int how often your platform has shown the post (drives the new-post push)
optional domain, language, is_ad, is_spoiler, upvote_ratio, awards, is_member

History item: post_id, ts_ms (when shown), visible_ms, focal_ms (attention), hover_ms, clicked, engaged, community, author, embedding. Set engaged to True for dwell β‰₯ 3 s or a click / like / save / join / reply. If you log only one dwell value, use it for both visible_ms and focal_ms, and set hover_ms to 0.

Output per candidate

  • score, position, topic_match;
  • v2b, cold_probability, tag_closeness;
  • warm_probability, warm_weight, engaged_history;
  • freshness_multiplier, push, has_content_embedding.

Log them to explain any feed.


6. Settings

HybridRecommender(**overrides); the defaults are stored in hybrid_engine/artifacts/meta.json.

Setting Default Effect
n0_engaged_posts 10 engaged posts at which the warm model reaches half of its maximum share
w_max 0.7 maximum warm-model share (the rest stays with the new-user model)
tag_bonus 0.1 weight of interest-tag closeness in the new-user score
decay_half_life_hours, decay_floor 720, 0.85 freshness: a week-old post keeps 98%, a month-old 93%, never below 85%
push_beta, push_max_age_hours 0.1, 24 lift for posts under 24 h old, fading with impressions
topic_first_slots, topic_block 7, 10 7 of every 10 positions go to posts matching the user's interests (0 = off)

Example: HybridRecommender(w_max=0.5, topic_first_slots=5).


7. Plug it into a backend

hybrid_engine/backend_adapter.py maps the data of a FastAPI + Redis + Postgres feed service onto the engine:

  • Posts: post_from_meta, post_from_row (backend posts β†’ candidates).
  • History: history_from_events (raw VIEW / LONG_VIEW / CLICK / LIKE / SAVE / SKIP / HIDE events β†’ history, merging visits and correcting timestamps).
  • Interests: expand_interests (broad interests β†’ tags).
  • Ranking: build_pool_hybrid, a drop-in replacement for a build_pool(r, uid, state, level) scoring step.
  • Vectors in Redis: encode_q3, fetch_q3.

Step-by-step guide: hybrid_engine/INTEGRATION.md. Speed: about 0.7 ms per candidate on one CPU thread, so a 200-post pool takes about 140 ms.


8. Continue training

pytorch_checkpoints/engine_v2.0/ holds the RecBole 1.2.0 originals of both models, 5 seeds each (cold_widedeep_seed0-4.pth, warm_deepfm_seed0-4.pth). The shipped ONNX files average these 5 seeds.

pip install torch recbole==1.2.0      # the checkpoints store a RecBole config, so RecBole must be importable to load them
import torch
ckpt = torch.load("pytorch_checkpoints/engine_v2.0/warm_deepfm_seed0.pth", map_location="cpu", weights_only=False)
print(ckpt["config"]["model"], ckpt["config"]["dataset"])          # DeepFM reco_ctr_nc_noid_strict
state = ckpt["state_dict"]                                         # weights

Rebuilding and retraining needs the RecBole datasets and scripts of the training kit: GitHub repository, branch recommender, folder models/training/. Use engine_v2/build_engine_v2.py, then engine_v2/verify_hybrid_engine.py. After retraining, re-export both ONNX files and meta.json together.


9. How it works

candidate posts ─┬─► tag closeness (8 tag-match features β†’ logistic score)
                 β”œβ”€β–Ί COLD model, Wide&Deep: post stats + content + tag match β†’ calibrated P(engage)
                 └─► WARM model, DeepFM: post inputs + 27 user-history features β†’ P(engage)

v2b   = P_cold + 0.1 Β· rank(tag closeness)                              (the new-user score)
w     = 0 for a user with no engaged posts, else min(0.7, engaged / (engaged + 10))
blend = (1 βˆ’ w) Β· rank(v2b) + w Β· rank(P_warm)
score = blend Β· (0.85 + 0.15 Β· 0.5^(age_h / 720))  +  0.1 Β· [age_h ≀ 24] / sqrt(1 + impressions)
order = in each block of 10 positions: 7 best interest matches, then 3 best of any kind
  • New users: w = 0, so pure v2b.
  • The warm model takes over gradually: w = 0.09 after 1 engaged post, 0.33 after 5, 0.50 after 10, and 0.70 at most. The new-user model always keeps at least 30%.
  • Excluded from the models: user and post IDs, feed position, scroll speed, and the community name and flair as direct inputs (the domain input still carries subreddit names for text posts; see the correction at the top).

10. Scores

All numbers are for v2.0, on data the models never trained on.

  • AUC: chance an engaged post is scored above a non-engaged one (0.5 = random).
  • NDCG@10: quality of the top 10 (1.0 = all engaged posts first).

Models (test split: last 15% of each user's timeline, 270 exposures, 44 engaged)

Model Validation AUC Test AUC
Cold model (Wide&Deep, 5-seed ensemble) 0.777 0.602
Warm model (DeepFM, 5-seed ensemble) 0.713 0.565

Ranking quality (NDCG@10; each feed ranked as one live request)

Feed v2.0 v2.0 with typical backend data same, no content vectors Random
Validation period, whole feeds 0.650 0.579 0.604 0.350
Validation period, 25-post pages 0.559 0.595 0.582 0.350
Test period, whole feeds 0.340 0.329 0.268 0.201
Test period, 25-post pages 0.415 0.397 0.441 0.286
Later real session 1, one feed 0.255 0.159 0.155 0.119
Later real session 1, 25-post pages 0.468 0.486 0.362 0.282
Later real session 2, one feed 0.110 0.110 0.110 0.079
Later real session 2, 25-post pages 0.403 0.382 0.372 0.245
Mean 0.400 0.380 0.362 0.239

"Typical backend data" means no flair, tags without confidences, a single dwell value and no vote ratio.

The two later real sessions (recorded two days after all training data)

Session 1 (89 posts, 11 engaged) Session 2 (142 posts, 10 engaged)
v2.0: NDCG@10 / engaged in top 10 / AUC 0.255 / 3 / 0.688 0.110 / 1 / 0.498
Reddit's own feed order 0.641 / 6 / 0.917 0.312 / 3 / 0.773
Random 0.124 / 1.1 / 0.50 0.070 / 0.6 / 0.50

New users (hybrid_engine/artifacts/new_user_report.json)

The same real feeds ranked as if each user had just signed up: their picked interests only, no history. NDCG@10:

Feed New user (interests only) New user, no interests picked Same user with full history Random
Validation period, whole feeds 0.525 0.564 0.650 0.350
Validation period, 25-post pages 0.525 0.568 0.559 0.350
Test period, whole feeds 0.410 0.303 0.340 0.201
Test period, 25-post pages 0.426 0.460 0.415 0.286
Later real session 1, one feed 0.221 0.445 0.255 0.119
Later real session 1, 25-post pages 0.341 0.527 0.468 0.282
Later real session 2, one feed 0.139 0.139 0.110 0.079
Later real session 2, 25-post pages 0.359 0.377 0.403 0.245
Mean 0.368 0.423 0.400 0.239
  • A brand-new user gets a useful feed from the first request: about 54% above random. The warm model is fully off for them (w = 0), as designed.
  • History improves results as it builds up (0.368 β†’ 0.400 on average).
  • A user who picks no interests is still served, ranked on post quality and content.
  • Caveat on "interests": these test users never picked interests; theirs were inferred from earlier behaviour, and some guesses were off (finance posts were pushed into the topic slots). That is why "no interests" scored highest here. Real signup picks should match users better. If your users' picks turn out to be unreliable, lower topic_first_slots.

Engineering checks

  • ONNX equals PyTorch within 2.4e-7.
  • The engine rebuilds the training features from raw data within 0.0021, with 0 category mismatches.
  • Engine scores match the offline predictions within 2.5e-4.
  • 9 automated tests pass.

Compared with the previous version: v1.1 used the community name and scored higher on its Reddit training data (test AUC 0.692 / 0.671, mean NDCG@10 0.446). Community names don't carry over to a new platform. Under realistic conditions the gap is small: 0.380 for v2.0 vs 0.410 for v1.1, and 0.362 vs 0.369 without content vectors. Details are in REPORT.md, section 13.


11. Model size

Model Parameters Shipped file
Cold model (Wide&Deep: 4-wide embeddings β†’ MLP 64 β†’ 32) 13,048 per seed cold_stage2.onnx, 5 seeds, 319 KB
Warm model (DeepFM: 16-wide embeddings β†’ MLP 64 β†’ 32) 66,457 per seed warm_deepfm.onnx, 5 seeds, 1.4 MB

The models are small on purpose: they learn from about 1,800 interactions, and a larger network would only memorise them. The .pth files are bigger than the parameter count because they also store the training configuration and optimizer state. The only large model in the pipeline is the public Qwen3-Embedding-0.6B text encoder (about 1.2 GB). It is used as is and downloaded from its own repository, so it is not copied here.


12. Where the engine's parts come from

This is an independent implementation built from methods used in well-known production recommenders. It is not affiliated with any of these companies.

Part Method Known use
Cold model Wide & Deep (Cheng et al., 2016) Google Play app recommendations (Google)
Warm model DeepFM (Guo et al., 2017) Huawei App Market CTR prediction (Huawei Noah's Ark Lab)
Cold β†’ warm blend new-user model handing over to a behavioural model as history grows the standard cold-start pattern of e-commerce and streaming feeds
Calibration Platt scaling CTR models in online advertising
Freshness decay exponential time decay the family of Hacker News and Reddit "hot" ranking
New-post push a boost for new items that fades with exposure initial-audience boosts for new content in short-video feeds such as TikTok
Topic-first slots + exploration share slot quotas, with 15–20% exploration kept in the serving layer quota-based feed blending in large feeds (YouTube, Instagram)
Content vectors Qwen3-Embedding-0.6B Alibaba Qwen's open embedding model
Interest tags LLM tagging into a closed taxonomy cold-starting item catalogs with LLM labels
Evaluated during research, not in v2.0 two-tower retrieval (YouTube), X's Thunder + Phoenix pipeline, MMR diversification, DCN V2 (Google), SASRec see REPORT.md

13. Training data and limitations

Data

  • Browser telemetry of real Reddit scrolling: 4 consenting participants, 1,794 post exposures, 1,677 posts. Usernames were hashed. The data is not published.
  • Label: engaged = click / vote / save / join / reply / opened post, or attention β‰₯ 3 s.
  • Split: chronological per user (70 / 15 / 15).
  • No look-ahead: history features only count interactions that ended before the ranked post was shown.

Limitations

  • Small data. Differences below about 0.1 AUC are within noise. Treat the settings as a starting point and re-tune w_max, n0_engaged_posts and decay/push on your own logs.
  • Trained on Reddit. Other platforms differ; retrain on your own data when you have it.
  • The topic-first rule can hide a user's new interests. Keep a 15–20% exploration share in your serving layer.
  • It ranks; it doesn't retrieve or moderate. Filter candidates for safety before ranking.

14. Archive (older versions)

Not for use. archive/ holds superseded research models, kept only so the history in REPORT.md can be checked. Everything in this README, and everything hybrid_engine/ loads, is v2.0. See archive/README.md.


@misc{unihive_recsys_2026,
  title  = {UniHive RecSys: a hybrid cold-to-warm feed-ranking engine},
  author = {Singh, Tirthanker},
  year   = {2026},
  url    = {https://huggingface.co/Tirthanker/UniHive_RecSys}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support