fix: use trafilatura.extract() for text + bare_extraction(with_metadata=True) for date/image ac179d1 skander101 commited on 22 days ago
fix: use BeautifulSoup for OG image extraction (handles name=, twitter:image, article:image, link[rel=image_src]) ca12da8 skander101 commited on 22 days ago
fix: walk ancestor chain, handle lazy-load attrs & relative URLs for images 76ae335 skander101 commited on 22 days ago
fix: extract dates from article pages via trafilatura instead of homepage HTML 623f05d skander101 commited on 22 days ago
Reformat SPONSOR_INFO, add owner_wikis, update templates/webapp 89ffd64 skander101 commited on 26 days ago
Replace World Health feeds: STAT, ScienceDaily, Science News, Medscape, NIH, Harvard Gazette 3c9e07f skander101 commited on 27 days ago
Fix sourcing penalty: detect Article A citing Article B from another outlet, not generic sourcing 1a87329 skander101 commited on 28 days ago
Add sourcing penalty: negative points for articles that source from other reporting 8914f3e skander101 commited on 28 days ago
Add sponsor modal per article with Wikipedia link, remove standalone sponsors page 769f921 skander101 commited on 28 days ago
Fix political leaning: expanded keywords, lower thresholds 8745ea5 skander101 commited on 28 days ago
Rich sponsor ownership: shareholders, types, ownership chains 84d1c3c skander101 commited on 28 days ago