Commit History
Merge pull request #2 from seanpedrick-case/dev 13d6f5f unverified
Sean Pedrick-Case commited on
Updated README.md to latest gradio version d95309d
Merge pull request #1 from seanpedrick-case/dev ccb541a unverified
Sean Pedrick-Case commited on
Fixes on representation model, visualisations, and embeddings in CPU mode. Package updates and optimisation for compatibility db3eaec
Changed llama-cpp version for cpu 39f4270
Corrected reference to sentence transformers dependency. Updated Dockerfile packages 611584f
Updated requirements beed391
Fix on returning GPU tensors to main function after embedding with zeroGPU. Representation model put under ZeroGPU spaces 02721f3
Removed explicit references to cuda in functions where spaces GPU are loaded cb7a4c9
zeroGPU spaces duration can now be defined from environment variable 38198b1
Added random seed to topic_core_funcs f957de1
Trying to load in cuda only within spaces environment to enable zero GPU space to run successfully 71afe01
Updated package versions in requirements files 5814ab0
Adjusted requirements for max available for Huggingface python==3.10 platform 6bf616b
Test update main requirements file for huggingface compatibility 9a4b420
Re-added TruncatedSVD dependency to topic_core_funcs.py f42e3d1
Debugged reference to random_seed in vectorisation and reference to torch in representation_model.py 8216d8c
Importing space package near start of app now to avoid issue with cuda being initialised before 9e84863
Llama-cpp-python in GPU mode doesn't seem to work well with Bertopic on Huggingface, so downgrading that to CPU version 88d81fa
Rearranged functions for embeddings creation to be compatible with zero GPU space. Updated packages. cc495e1
Added and replaced relevant files to download in download_model.py to allow for app use on AWS 49e0db8
Updated Dockerfile with latest packages 08eb30d
Added example of how to run function from command line. Updated packages. Embedding model default now smaller and at fp16. 34f1e83
Improved initial clean options. Now has option to return embeddings only. 89c4d20
Corrected minor Dockerfile package version issue 593153e
App now retains original index following cleaning to allow for referring back to original data 90553eb
Now installed dependencies into correct folder in Dockerfile 5888649
Finally managed to enforce cpu torch install in Dockerfile 97913c4
Further optimised Dockerfile and requirements (smaller torch installation now hopefully) 00db72b
Transferring across installed packages from build stage in Dockerfile c9da99d
Changed Dockerfile to multi-stage build to further reduce size 0fd155c
Trying to make container image smaller through Dockerfile 7d5387e
Minor changes to reduce Dockerfile size b767539
Updated download_model.py to download pytorch .bin file 1c0bfd4
Removed some requirements from Dockerfile for AWS deployment to reduce container size 51ba1cb
Added NUMBA_CACHE_DIR to Docker environmental variables cd6a3e0
Allowed for app running on AWS to use smaller embedding model and not to load representation LLM (due to size restrictions). 22ca76e
Dockerfile now installs models directly into user folder instead of moving from base folder 3c1c3de
Updated Gradio version for spaces. Updated Dockerfile to enable Llama.cpp build with Cmake d34af22
Only aggregate topics not 'other', allowed for minimum sentence length, default max_topics now will auto aggregate topics. Added Cognito Auth functionality (boto3 with AWS). 1e2bb3e
Can split passages into sentences. Improved embedding, LLM representation models, improved zero shot capabilities 55f0ce3
Updated packages. Improve hierarchy vis. Better models - mixedbread and phi3. Now option to split texts into sentences before modelling. 04a15c5
Minor cleaning, csv formatting changes d80c8f5
Sean-Case commited on
Reduce outliers now more efficient and relabels with correct vectoriser. Default topic labels now tidier. Hiearchical topics outputs more useful for joining to df afterwards. Switched low resource reduction algorithm to UMAP as default is not good. e1c1f68
Should now parse custom regex correctly. Will now wipe previously created embeddings if 'low resource mode' option switched. 0a543a0
Sean-Case commited on
Allowed for uploading custom regex for cleaning. Fixed calculate all probabilities, reduce outliers. Added text tree for hierarchical modelling. 381f959
Upgraded to Gradio 4.16.0. Guide for converting to exe added. 0a177ca
Hopefully now LLM download from hub should work cdcd7af
Note about LLM not working now successfully added! e2dfc1e
Sean-Case commited on