Text Generation
Transformers
Safetensors
French
English
Chinese
deepseek_v4
cortex
code-generation
web-development
software-engineering
Mixture of Experts
8-bit precision
fp8
Instructions to use Frankenstein-Labs/cortex.6.sol with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Frankenstein-Labs/cortex.6.sol with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Frankenstein-Labs/cortex.6.sol")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Frankenstein-Labs/cortex.6.sol") model = AutoModelForCausalLM.from_pretrained("Frankenstein-Labs/cortex.6.sol", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Frankenstein-Labs/cortex.6.sol with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Frankenstein-Labs/cortex.6.sol" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Frankenstein-Labs/cortex.6.sol
- SGLang
How to use Frankenstein-Labs/cortex.6.sol with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/cortex.6.sol" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/cortex.6.sol" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Frankenstein-Labs/cortex.6.sol with Docker Model Runner:
docker model run hf.co/Frankenstein-Labs/cortex.6.sol
Download cortex_ai/client.py from Frankenstein-Labs/cortex.6.sol: direct link, hf CLI and curl.
- Browser
- Download file 3.27 kB
-
https://huggingface.co/Frankenstein-Labs/cortex.6.sol/resolve/main/cortex_ai/client.py
- Command line
-
hf download hf://Frankenstein-Labs/cortex.6.sol/cortex_ai/client.py
-
curl -L -o client.py https://huggingface.co/Frankenstein-Labs/cortex.6.sol/resolve/main/cortex_ai/client.py
3.27 kB
| """Minimal Python client for the CORTEX AI API. | |
| from cortex_ai.client import CortexClient | |
| client = CortexClient("http://localhost:8000") | |
| print(client.ask("Combien font 12 * 8 ?").content) | |
| """ | |
| from __future__ import annotations | |
| import json | |
| import urllib.error | |
| import urllib.request | |
| from dataclasses import dataclass, field | |
| from typing import Any | |
| class Reply: | |
| """One assistant reply, with the reasoning trace and any tool calls.""" | |
| content: str | |
| reasoning: str = "" | |
| tool_calls: list[dict[str, Any]] = field(default_factory=list) | |
| usage: dict[str, int] = field(default_factory=dict) | |
| raw: dict[str, Any] = field(default_factory=dict) | |
| class CortexError(Exception): | |
| """Raised when the API returns an error.""" | |
| class CortexClient: | |
| """Talks to a running CORTEX AI server over HTTP.""" | |
| def __init__(self, base_url: str = "http://localhost:8000", api_key: str = "") -> None: | |
| self.base_url = base_url.rstrip("/") | |
| self.api_key = api_key | |
| self._history: list[dict[str, str]] = [] | |
| def _post(self, path: str, payload: dict[str, Any]) -> dict[str, Any]: | |
| request = urllib.request.Request( | |
| self.base_url + path, | |
| data=json.dumps(payload).encode("utf-8"), | |
| headers={"Content-Type": "application/json"}, | |
| method="POST", | |
| ) | |
| if self.api_key: | |
| request.add_header("Authorization", f"Bearer {self.api_key}") | |
| try: | |
| with urllib.request.urlopen(request, timeout=600) as response: | |
| return json.loads(response.read().decode("utf-8")) | |
| except urllib.error.HTTPError as exc: | |
| raise CortexError(f"HTTP {exc.code}: {exc.read().decode('utf-8', 'replace')}") from None | |
| except urllib.error.URLError as exc: | |
| raise CortexError(f"cannot reach {self.base_url}: {exc.reason}") from None | |
| def _get(self, path: str) -> dict[str, Any]: | |
| request = urllib.request.Request(self.base_url + path) | |
| if self.api_key: | |
| request.add_header("Authorization", f"Bearer {self.api_key}") | |
| with urllib.request.urlopen(request, timeout=60) as response: | |
| return json.loads(response.read().decode("utf-8")) | |
| def health(self) -> dict[str, Any]: | |
| return self._get("/health") | |
| def models(self) -> list[str]: | |
| return [m["id"] for m in self._get("/v1/models")["data"]] | |
| def ask(self, message: str, *, keep_history: bool = False) -> Reply: | |
| """Send one user message and return the reply.""" | |
| messages = [*self._history, {"role": "user", "content": message}] | |
| body = self._post("/v1/chat/completions", {"messages": messages}) | |
| choice = body["choices"][0]["message"] | |
| reply = Reply( | |
| content=choice["content"], | |
| reasoning=body.get("reasoning_content", ""), | |
| tool_calls=body.get("tool_calls", []), | |
| usage=body.get("usage", {}), | |
| raw=body, | |
| ) | |
| if keep_history: | |
| self._history = [ | |
| *messages, | |
| {"role": "assistant", "content": reply.content}, | |
| ] | |
| return reply | |
| def reset(self) -> None: | |
| """Forget the conversation history.""" | |
| self._history = [] | |