- UniHive RecSys: hybrid feed-ranking engine (v2.0)
- Contents
- 1. Which files you need
- 2. Install and run in 2 minutes
- 3. Use it in your code
- 4. Content vectors (recommended)
- 5. Inputs and outputs
- 6. Settings
- 7. Plug it into a backend
- 8. Continue training
- 9. How it works
- 10. Scores
- 11. Model size
- 12. Where the engine's parts come from
- 13. Training data and limitations
- 14. Archive (older versions)
- Contents
UniHive RecSys: hybrid feed-ranking engine (v2.0)
A small, fast post recommendation engine for a student social app. Give it a user (interests + past interactions) and a list of candidate posts; it returns the posts in feed order, with a score and an explanation for each.
β This repository is engine v2.0: our final, intended design. Use
hybrid_engine/.
- Run it:
hybrid_engine/(code) +hybrid_engine/artifacts/(the weights it loads). Nothing else is needed.- Train it further:
pytorch_checkpoints/engine_v2.0/(the same v2.0 models as PyTorch files).- Ignore
archive/unless you are studying the project's history. Those are superseded experiments, built on different inputs, and they will not work withhybrid_engine/.
Correction (2026-10-07). Version 2.0 removes the community name and flair as direct inputs and from the embedding text. It still receives each post's
domainfield, which for Reddit text posts isself.<subreddit>: 24 of the 30 domain tokens known to both models are subreddit names. Community information therefore remains indirectly available to the 2.0 models for text posts. A follow-up release removesdomainentirely. Until then, treat v2.0's community independence as partial.
Contents
- Which files you need
- Install and run in 2 minutes
- Use it in your code
- Content vectors (recommended)
- Inputs and outputs
- Settings
- Plug it into a backend
- Continue training
- How it works
- Scores
- Model size
- Where the engine's parts come from
- Training data and limitations
- Archive (older versions)
1. Which files you need
| You want to... | Files |
|---|---|
| Rank posts (most users) | hybrid_engine/ (code + artifacts/cold_stage2.onnx, artifacts/warm_deepfm.onnx, artifacts/meta.json) |
| Turn post text into content vectors | the public encoder Qwen/Qwen3-Embedding-0.6B (section 4) |
| Continue training the models | pytorch_checkpoints/engine_v2.0/ (section 8) |
| Read the full project report | REPORT.md |
The engine reads exactly the files in hybrid_engine/artifacts/. Don't replace them with files from archive/: the feature lists, vocabularies and standardisation differ.
2. Install and run in 2 minutes
pip install huggingface_hub numpy onnxruntime
hf download Tirthanker/UniHive_RecSys --include "hybrid_engine/*" --local-dir UniHive_RecSys
cd UniHive_RecSys
python -m hybrid_engine.example # demo: a brand-new user vs the same user after 30 engaged posts
python -m hybrid_engine.tests.test_engine # smoke tests; should end with "community / flair guard passed"
Python 3.10+. Runs on CPU; no GPU needed.
3. Use it in your code
import time
from hybrid_engine import HybridRecommender
engine = HybridRecommender() # load once at startup; reuse for every request
user = {
"interests": ["machine-learning", "computer-vision"], # tag slugs the user picked (list: meta.json -> taxonomy.child)
"history": [ # what the user saw before; may be empty for a new user
{"post_id": "p1", "ts_ms": 1791141039157, "visible_ms": 5200, "focal_ms": 4800, "hover_ms": 300,
"clicked": True, "engaged": True, "community": "c-ml", "author": "a-17", "embedding": None},
],
}
candidates = [
{"post_id": "p9", "post_type": "text", "score": 42, "comments": 7, "age_hours": 3.5,
"title": "Fine-tuning ViTs on a laptop GPU", "body": "Notes from a weekend project ...",
"tags": {"computer-vision": 0.9, "deep-learning": 0.7}, "embedding": None,
"author": "a-3", "impressions": 0, "community": "c-cv"},
{"post_id": "p10", "post_type": "image", "score": 300, "comments": 40, "age_hours": 50,
"title": "Campus fest photos", "body": "", "tags": {"photography": 0.8}, "embedding": None,
"author": "a-8", "impressions": 120, "community": "c-fest"},
]
feed = engine.recommend(user, candidates, now_ms=time.time() * 1000, k=10)
for d in feed:
print(d["position"], d["post_id"], round(d["score"], 3), "topic" if d["topic_match"] else "", "warm weight", d["warm_weight"])
recommend(user, candidates, now_ms, k): the top k in final feed order.rank(...): every candidate in feed order.score(...): raw scores without the topic-first reordering.
now_ms is the request time in milliseconds. History items at or after it are ignored automatically.
4. Content vectors (recommended)
The models read a 1024-d content vector per post. Build it exactly like this (the same recipe the models were trained on):
from sentence_transformers import SentenceTransformer # pip install sentence-transformers
from hybrid_engine.backend_adapter import qwen3_text
encoder = SentenceTransformer("Qwen/Qwen3-Embedding-0.6B"); encoder.max_seq_length = 512 # load once
text = qwen3_text(title, body, image_caption) # "title. body" (+ " Image: caption"); no community, no flair
post["embedding"] = encoder.encode(text, normalize_embeddings=True)
- Compute it once per post when the post is created, and store it.
- Without vectors (
"embedding": None) the engine still ranks, but loses the content signal: about 0.02 lower NDCG@10 on average (section 10). - A vector of the wrong size (e.g. a 512-d CLIP vector) is ignored on purpose: it lives in a different space.
5. Inputs and outputs
Candidate post
| Field | Type | Meaning |
|---|---|---|
post_id |
str | your id |
post_type |
str | text, image, video, link, gallery, ... |
score, comments |
int | likes / upvotes and comment count |
age_hours |
float | hours since the post was created |
title, body, image_caption |
str | used for length features (and for the content vector, section 4) |
tags |
dict or list | {slug: confidence} from the 114-tag taxonomy, or a list of slugs (confidence 1.0) |
embedding |
1024 floats or None | section 4 |
author, community |
str | used only for the user's own history with that author / community, never as model inputs |
impressions |
int | how often your platform has shown the post (drives the new-post push) |
| optional | domain, language, is_ad, is_spoiler, upvote_ratio, awards, is_member |
History item: post_id, ts_ms (when shown), visible_ms, focal_ms (attention), hover_ms, clicked, engaged, community, author, embedding. Set engaged to True for dwell β₯ 3 s or a click / like / save / join / reply. If you log only one dwell value, use it for both visible_ms and focal_ms, and set hover_ms to 0.
Output per candidate
score,position,topic_match;v2b,cold_probability,tag_closeness;warm_probability,warm_weight,engaged_history;freshness_multiplier,push,has_content_embedding.
Log them to explain any feed.
6. Settings
HybridRecommender(**overrides); the defaults are stored in hybrid_engine/artifacts/meta.json.
| Setting | Default | Effect |
|---|---|---|
n0_engaged_posts |
10 | engaged posts at which the warm model reaches half of its maximum share |
w_max |
0.7 | maximum warm-model share (the rest stays with the new-user model) |
tag_bonus |
0.1 | weight of interest-tag closeness in the new-user score |
decay_half_life_hours, decay_floor |
720, 0.85 | freshness: a week-old post keeps 98%, a month-old 93%, never below 85% |
push_beta, push_max_age_hours |
0.1, 24 | lift for posts under 24 h old, fading with impressions |
topic_first_slots, topic_block |
7, 10 | 7 of every 10 positions go to posts matching the user's interests (0 = off) |
Example: HybridRecommender(w_max=0.5, topic_first_slots=5).
7. Plug it into a backend
hybrid_engine/backend_adapter.py maps the data of a FastAPI + Redis + Postgres feed service onto the engine:
- Posts:
post_from_meta,post_from_row(backend posts β candidates). - History:
history_from_events(raw VIEW / LONG_VIEW / CLICK / LIKE / SAVE / SKIP / HIDE events β history, merging visits and correcting timestamps). - Interests:
expand_interests(broad interests β tags). - Ranking:
build_pool_hybrid, a drop-in replacement for abuild_pool(r, uid, state, level)scoring step. - Vectors in Redis:
encode_q3,fetch_q3.
Step-by-step guide: hybrid_engine/INTEGRATION.md. Speed: about 0.7 ms per candidate on one CPU thread, so a 200-post pool takes about 140 ms.
8. Continue training
pytorch_checkpoints/engine_v2.0/ holds the RecBole 1.2.0 originals of both models, 5 seeds each (cold_widedeep_seed0-4.pth, warm_deepfm_seed0-4.pth). The shipped ONNX files average these 5 seeds.
pip install torch recbole==1.2.0 # the checkpoints store a RecBole config, so RecBole must be importable to load them
import torch
ckpt = torch.load("pytorch_checkpoints/engine_v2.0/warm_deepfm_seed0.pth", map_location="cpu", weights_only=False)
print(ckpt["config"]["model"], ckpt["config"]["dataset"]) # DeepFM reco_ctr_nc_noid_strict
state = ckpt["state_dict"] # weights
Rebuilding and retraining needs the RecBole datasets and scripts of the training kit: GitHub repository, branch recommender, folder models/training/. Use engine_v2/build_engine_v2.py, then engine_v2/verify_hybrid_engine.py. After retraining, re-export both ONNX files and meta.json together.
9. How it works
candidate posts ββ¬ββΊ tag closeness (8 tag-match features β logistic score)
βββΊ COLD model, Wide&Deep: post stats + content + tag match β calibrated P(engage)
βββΊ WARM model, DeepFM: post inputs + 27 user-history features β P(engage)
v2b = P_cold + 0.1 Β· rank(tag closeness) (the new-user score)
w = 0 for a user with no engaged posts, else min(0.7, engaged / (engaged + 10))
blend = (1 β w) Β· rank(v2b) + w Β· rank(P_warm)
score = blend Β· (0.85 + 0.15 Β· 0.5^(age_h / 720)) + 0.1 Β· [age_h β€ 24] / sqrt(1 + impressions)
order = in each block of 10 positions: 7 best interest matches, then 3 best of any kind
- New users: w = 0, so pure v2b.
- The warm model takes over gradually: w = 0.09 after 1 engaged post, 0.33 after 5, 0.50 after 10, and 0.70 at most. The new-user model always keeps at least 30%.
- Excluded from the models: user and post IDs, feed position, scroll speed, and the community name and flair as direct inputs (the
domaininput still carries subreddit names for text posts; see the correction at the top).
10. Scores
All numbers are for v2.0, on data the models never trained on.
- AUC: chance an engaged post is scored above a non-engaged one (0.5 = random).
- NDCG@10: quality of the top 10 (1.0 = all engaged posts first).
Models (test split: last 15% of each user's timeline, 270 exposures, 44 engaged)
| Model | Validation AUC | Test AUC |
|---|---|---|
| Cold model (Wide&Deep, 5-seed ensemble) | 0.777 | 0.602 |
| Warm model (DeepFM, 5-seed ensemble) | 0.713 | 0.565 |
Ranking quality (NDCG@10; each feed ranked as one live request)
| Feed | v2.0 | v2.0 with typical backend data | same, no content vectors | Random |
|---|---|---|---|---|
| Validation period, whole feeds | 0.650 | 0.579 | 0.604 | 0.350 |
| Validation period, 25-post pages | 0.559 | 0.595 | 0.582 | 0.350 |
| Test period, whole feeds | 0.340 | 0.329 | 0.268 | 0.201 |
| Test period, 25-post pages | 0.415 | 0.397 | 0.441 | 0.286 |
| Later real session 1, one feed | 0.255 | 0.159 | 0.155 | 0.119 |
| Later real session 1, 25-post pages | 0.468 | 0.486 | 0.362 | 0.282 |
| Later real session 2, one feed | 0.110 | 0.110 | 0.110 | 0.079 |
| Later real session 2, 25-post pages | 0.403 | 0.382 | 0.372 | 0.245 |
| Mean | 0.400 | 0.380 | 0.362 | 0.239 |
"Typical backend data" means no flair, tags without confidences, a single dwell value and no vote ratio.
The two later real sessions (recorded two days after all training data)
| Session 1 (89 posts, 11 engaged) | Session 2 (142 posts, 10 engaged) | |
|---|---|---|
| v2.0: NDCG@10 / engaged in top 10 / AUC | 0.255 / 3 / 0.688 | 0.110 / 1 / 0.498 |
| Reddit's own feed order | 0.641 / 6 / 0.917 | 0.312 / 3 / 0.773 |
| Random | 0.124 / 1.1 / 0.50 | 0.070 / 0.6 / 0.50 |
New users (hybrid_engine/artifacts/new_user_report.json)
The same real feeds ranked as if each user had just signed up: their picked interests only, no history. NDCG@10:
| Feed | New user (interests only) | New user, no interests picked | Same user with full history | Random |
|---|---|---|---|---|
| Validation period, whole feeds | 0.525 | 0.564 | 0.650 | 0.350 |
| Validation period, 25-post pages | 0.525 | 0.568 | 0.559 | 0.350 |
| Test period, whole feeds | 0.410 | 0.303 | 0.340 | 0.201 |
| Test period, 25-post pages | 0.426 | 0.460 | 0.415 | 0.286 |
| Later real session 1, one feed | 0.221 | 0.445 | 0.255 | 0.119 |
| Later real session 1, 25-post pages | 0.341 | 0.527 | 0.468 | 0.282 |
| Later real session 2, one feed | 0.139 | 0.139 | 0.110 | 0.079 |
| Later real session 2, 25-post pages | 0.359 | 0.377 | 0.403 | 0.245 |
| Mean | 0.368 | 0.423 | 0.400 | 0.239 |
- A brand-new user gets a useful feed from the first request: about 54% above random. The warm model is fully off for them (w = 0), as designed.
- History improves results as it builds up (0.368 β 0.400 on average).
- A user who picks no interests is still served, ranked on post quality and content.
- Caveat on "interests": these test users never picked interests; theirs were inferred from earlier behaviour, and some guesses were off (finance posts were pushed into the topic slots). That is why "no interests" scored highest here. Real signup picks should match users better. If your users' picks turn out to be unreliable, lower
topic_first_slots.
Engineering checks
- ONNX equals PyTorch within 2.4e-7.
- The engine rebuilds the training features from raw data within 0.0021, with 0 category mismatches.
- Engine scores match the offline predictions within 2.5e-4.
- 9 automated tests pass.
Compared with the previous version: v1.1 used the community name and scored higher on its Reddit training data (test AUC 0.692 / 0.671, mean NDCG@10 0.446). Community names don't carry over to a new platform. Under realistic conditions the gap is small: 0.380 for v2.0 vs 0.410 for v1.1, and 0.362 vs 0.369 without content vectors. Details are in REPORT.md, section 13.
11. Model size
| Model | Parameters | Shipped file |
|---|---|---|
| Cold model (Wide&Deep: 4-wide embeddings β MLP 64 β 32) | 13,048 per seed | cold_stage2.onnx, 5 seeds, 319 KB |
| Warm model (DeepFM: 16-wide embeddings β MLP 64 β 32) | 66,457 per seed | warm_deepfm.onnx, 5 seeds, 1.4 MB |
The models are small on purpose: they learn from about 1,800 interactions, and a larger network would only memorise them. The .pth files are bigger than the parameter count because they also store the training configuration and optimizer state. The only large model in the pipeline is the public Qwen3-Embedding-0.6B text encoder (about 1.2 GB). It is used as is and downloaded from its own repository, so it is not copied here.
12. Where the engine's parts come from
This is an independent implementation built from methods used in well-known production recommenders. It is not affiliated with any of these companies.
| Part | Method | Known use |
|---|---|---|
| Cold model | Wide & Deep (Cheng et al., 2016) | Google Play app recommendations (Google) |
| Warm model | DeepFM (Guo et al., 2017) | Huawei App Market CTR prediction (Huawei Noah's Ark Lab) |
| Cold β warm blend | new-user model handing over to a behavioural model as history grows | the standard cold-start pattern of e-commerce and streaming feeds |
| Calibration | Platt scaling | CTR models in online advertising |
| Freshness decay | exponential time decay | the family of Hacker News and Reddit "hot" ranking |
| New-post push | a boost for new items that fades with exposure | initial-audience boosts for new content in short-video feeds such as TikTok |
| Topic-first slots + exploration share | slot quotas, with 15β20% exploration kept in the serving layer | quota-based feed blending in large feeds (YouTube, Instagram) |
| Content vectors | Qwen3-Embedding-0.6B | Alibaba Qwen's open embedding model |
| Interest tags | LLM tagging into a closed taxonomy | cold-starting item catalogs with LLM labels |
| Evaluated during research, not in v2.0 | two-tower retrieval (YouTube), X's Thunder + Phoenix pipeline, MMR diversification, DCN V2 (Google), SASRec | see REPORT.md |
13. Training data and limitations
Data
- Browser telemetry of real Reddit scrolling: 4 consenting participants, 1,794 post exposures, 1,677 posts. Usernames were hashed. The data is not published.
- Label: engaged = click / vote / save / join / reply / opened post, or attention β₯ 3 s.
- Split: chronological per user (70 / 15 / 15).
- No look-ahead: history features only count interactions that ended before the ranked post was shown.
Limitations
- Small data. Differences below about 0.1 AUC are within noise. Treat the settings as a starting point and re-tune
w_max,n0_engaged_postsand decay/push on your own logs. - Trained on Reddit. Other platforms differ; retrain on your own data when you have it.
- The topic-first rule can hide a user's new interests. Keep a 15β20% exploration share in your serving layer.
- It ranks; it doesn't retrieve or moderate. Filter candidates for safety before ranking.
14. Archive (older versions)
Not for use. archive/ holds superseded research models, kept only so the history in REPORT.md can be checked. Everything in this README, and everything hybrid_engine/ loads, is v2.0. See archive/README.md.
@misc{unihive_recsys_2026,
title = {UniHive RecSys: a hybrid cold-to-warm feed-ranking engine},
author = {Singh, Tirthanker},
year = {2026},
url = {https://huggingface.co/Tirthanker/UniHive_RecSys}
}