Spaces:
Running
Running
Download LLM_Basics.html from rr19tech/RetroPosterBoard: direct link, hf CLI and curl.
- Browser
- Download file 4.18 kB
-
https://huggingface.co/spaces/rr19tech/RetroPosterBoard/resolve/main/LLM_Basics.html
- Command line
-
hf download hf://spaces/rr19tech/RetroPosterBoard/LLM_Basics.html
-
curl -L -o LLM_Basics.html https://huggingface.co/spaces/rr19tech/RetroPosterBoard/resolve/main/LLM_Basics.html
4.18 kB
| <html lang="en"> | |
| <head> | |
| <meta charset="UTF-8"> | |
| <meta name="viewport" content="width=device-width, initial-scale=1.0"> | |
| <title>A Graph RAG study - Experimental Setup</title> | |
| <style> | |
| body { | |
| font-family: Arial, sans-serif; | |
| line-height: 1.6; | |
| margin: 20px; | |
| } | |
| h1, h2 { | |
| color: #333; | |
| } | |
| h2 { | |
| margin-top: 30px; | |
| } | |
| ul { | |
| list-style-type: disc; | |
| margin-left: 20px; | |
| } | |
| p { | |
| margin-bottom: 15px; | |
| } | |
| table{ | |
| border-collapse: collapse; | |
| width: 95%; | |
| border: 2px solid #2c3e50; | |
| } | |
| tr{ | |
| border-bottom: 2px solid #b60e0e; | |
| } | |
| td{ | |
| width: 15%; | |
| vertical-align: top; | |
| border: 2px solid #3498db; | |
| } | |
| </style> | |
| </head> | |
| <body> | |
| <h2>The very basics</h2> | |
| <p> | |
| LLM model binary files irrespective of the format GGUF, SafeTensors, Onnx, Pytorch etc, the | |
| file only contains raw weights, metadata, and tokenizers, not the code required to execute mathematical operations. | |
| You need inference engines like llama.cpp, which will map these weights to CPU/GPU etc.<br> | |
| llamafile combines these engine and gguf file into a single binary so we can run it. | |
| Some of the other popular options to run models locally include ollama, mlx_lm, LM Studio, vLLM etc. | |
| </p> | |
| <h2>So, now we got the model locally, run the engine, so what next</h2> | |
| <p> | |
| We may run the interactive chat window that often these engines provide to chat or for generative tasks.<br> | |
| But more often we need to interact with these models programatically thru API's so we can make applications utilize the LLM capabilities.<br> | |
| Examine the following code from langchain to interact with openai based models | |
| <pre> | |
| from langchain_openai import ChatOpenAI | |
| | | |
| model = ChatOpenAI( | |
| model="...", | |
| temperature=0, | |
| max_tokens=None, | |
| timeout=None, | |
| max_retries=2, | |
| # api_key="...", | |
| # base_url="...", | |
| # organization="...", | |
| # other params... | |
| ) | |
| </pre> | |
| Or a native OpenAI, few liner to chat with a local model and get an answer. | |
| <pre> | |
| from openai import OpenAI | |
| client = OpenAI(base_url="http://HOST:PORT/v1", api_key="None") | |
| response = client.chat.completions.create( | |
| model="...", | |
| messages=[ | |
| {"role": "user", "content": QUESTION} | |
| ] | |
| ) | |
| print(response.choices[0].message.content) | |
| </pre> | |
| Please go through the following hyper-parameters, that can be set externally, that can control the behaviour of the LLM. | |
| <pre> | |
| model / model_name: Name of the OpenAI model to use (e.g., gpt-4o). | |
| temperature: Sampling randomness controlling creativity (float). | |
| max_tokens: Maximum number of tokens to generate in the completion. | |
| max_completion_tokens: Maximum upper bound for reasoning and output tokens for newer reasoning models. | |
| seed: Integer random seed for deterministic sampling. | |
| frequency_penalty: Penalizes repeated tokens based on their frequency. | |
| presence_penalty: Penalizes repeated tokens to encourage new topics. | |
| logprobs: Boolean whether to return log probabilities of output tokens. | |
| top_logprobs: Number of most likely tokens to return log probabilities for at each position. | |
| </pre> | |
| </p> | |
| </body> | |
| </html> |