Text Generation
Transformers
Safetensors
French
English
Chinese
deepseek_v4
cortex
code-generation
web-development
software-engineering
Mixture of Experts
8-bit precision
fp8
Instructions to use Frankenstein-Labs/cortex.6.sol with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Frankenstein-Labs/cortex.6.sol with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Frankenstein-Labs/cortex.6.sol")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Frankenstein-Labs/cortex.6.sol") model = AutoModelForCausalLM.from_pretrained("Frankenstein-Labs/cortex.6.sol", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Frankenstein-Labs/cortex.6.sol with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Frankenstein-Labs/cortex.6.sol" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Frankenstein-Labs/cortex.6.sol
- SGLang
How to use Frankenstein-Labs/cortex.6.sol with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/cortex.6.sol" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/cortex.6.sol" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Frankenstein-Labs/cortex.6.sol with Docker Model Runner:
docker model run hf.co/Frankenstein-Labs/cortex.6.sol
|
Download docs/api.fr.md from Frankenstein-Labs/cortex.6.sol: direct link, hf CLI and curl.
- Browser
- Download file 5.61 kB
-
https://huggingface.co/Frankenstein-Labs/cortex.6.sol/resolve/main/docs/api.fr.md
- Command line
-
hf download hf://Frankenstein-Labs/cortex.6.sol/docs/api.fr.md
-
curl -L -o api.fr.md https://huggingface.co/Frankenstein-Labs/cortex.6.sol/resolve/main/docs/api.fr.md
5.61 kB
| # L'API CORTEX AI | |
| CORTEX AI expose une API **compatible OpenAI**. Tout client qui parle ce | |
| protocole fonctionne en changeant seulement l'URL de base : le SDK Python | |
| officiel, LangChain, LlamaIndex, ou un simple `curl`. | |
| --- | |
| ## Lancer le serveur | |
| ```bash | |
| pip install -r requirements-ecosystem.txt | |
| export PYTHONPATH=/chemin/vers/Cortex-ai:/chemin/vers/Cortex-ai/encoding | |
| python -m cortex_ai.api.serve | |
| ``` | |
| Par défaut, le serveur démarre avec l'adaptateur simulé : aucune carte | |
| graphique, aucun téléchargement. Pour charger le vrai modèle : | |
| ```bash | |
| export CORTEX_ADAPTER=hf | |
| ``` | |
| --- | |
| ## Configuration | |
| Toutes les options passent par des variables d'environnement. | |
| | Variable | Défaut | Effet | | |
| |---|---|---| | |
| | `CORTEX_MODEL_ID` | `Frankenstein-Labs/Cortex-ai` | Identifiant du modèle | | |
| | `CORTEX_HOST` | `0.0.0.0` | Adresse d'écoute | | |
| | `CORTEX_PORT` | `8000` | Port | | |
| | `CORTEX_API_KEY` | vide | Si défini, authentification obligatoire | | |
| | `CORTEX_THINKING_MODE` | `thinking` | `chat` ou `thinking` | | |
| | `CORTEX_REASONING_EFFORT` | `high` | `low`, `high` ou `max` | | |
| | `CORTEX_MAX_TOOL_ROUNDS` | `8` | Nombre maximal de tours d'outils | | |
| | `CORTEX_TOOLS` | toutes | Liste séparée par des virgules | | |
| | `CORTEX_ADAPTER` | `mock` | `mock` ou `hf` | | |
| --- | |
| ## Points d'accès | |
| ### `GET /health` | |
| Vérifie que le serveur répond. | |
| ```bash | |
| curl http://localhost:8000/health | |
| ``` | |
| ```json | |
| { | |
| "status": "ok", | |
| "model": "Frankenstein-Labs/Cortex-ai", | |
| "tools": ["calculate", "cortex_identity", "current_time", "text_stats"], | |
| "thinking_mode": "thinking" | |
| } | |
| ``` | |
| ### `GET /v1/models` | |
| Liste les modèles disponibles, au format OpenAI. | |
| ```bash | |
| curl http://localhost:8000/v1/models | |
| ``` | |
| ### `POST /v1/chat/completions` | |
| Le point d'accès principal. | |
| ```bash | |
| curl http://localhost:8000/v1/chat/completions \ | |
| -H "Content-Type: application/json" \ | |
| -d '{ | |
| "model": "Frankenstein-Labs/Cortex-ai", | |
| "messages": [{"role": "user", "content": "Combien font 12 * 8 ?"}] | |
| }' | |
| ``` | |
| Réponse : | |
| ```json | |
| { | |
| "id": "chatcmpl-6bd0cefc9196421ba09b74d9", | |
| "object": "chat.completion", | |
| "created": 1789761504, | |
| "model": "Frankenstein-Labs/Cortex-ai", | |
| "choices": [ | |
| { | |
| "index": 0, | |
| "message": {"role": "assistant", "content": "..."}, | |
| "finish_reason": "stop" | |
| } | |
| ], | |
| "usage": {"prompt_tokens": 6, "completion_tokens": 9, "total_tokens": 15}, | |
| "reasoning_content": "J'ai reçu le résultat de l'outil, je peux conclure.", | |
| "tool_calls": [ | |
| {"name": "calculate", "arguments": {"expression": "12 * 8"}, "result": "96", "ok": true} | |
| ] | |
| } | |
| ``` | |
| Deux champs s'ajoutent au format OpenAI : | |
| | Champ | Contenu | | |
| |---|---| | |
| | `reasoning_content` | Le raisonnement du modèle, en mode `thinking` | | |
| | `tool_calls` | Les outils réellement exécutés, avec leur résultat | | |
| Ces champs sont additifs : un client OpenAI standard les ignore sans erreur. | |
| --- | |
| ## Le SDK OpenAI | |
| Le SDK officiel fonctionne tel quel. | |
| ```python | |
| from openai import OpenAI | |
| client = OpenAI(base_url="http://localhost:8000/v1", api_key="peu-importe") | |
| reponse = client.chat.completions.create( | |
| model="Frankenstein-Labs/Cortex-ai", | |
| messages=[{"role": "user", "content": "Combien font 12 * 8 ?"}], | |
| ) | |
| print(reponse.choices[0].message.content) | |
| ``` | |
| Si vous avez défini `CORTEX_API_KEY`, passez la même valeur dans `api_key`. | |
| --- | |
| ## Le client Python inclus | |
| Un client minimal, sans dépendance, est fourni. | |
| ```python | |
| from cortex_ai.client import CortexClient | |
| client = CortexClient("http://localhost:8000") | |
| print(client.health()["status"]) | |
| print(client.models()) | |
| reponse = client.ask("Combien font 12 * 8 ?") | |
| print(reponse.content) # la réponse | |
| print(reponse.reasoning) # le raisonnement | |
| print(reponse.tool_calls) # les outils exécutés | |
| print(reponse.usage) # les compteurs de jetons | |
| ``` | |
| Pour conserver l'historique entre les appels : | |
| ```python | |
| client.ask("Je m'appelle Abdoulaye.", keep_history=True) | |
| client.ask("Comment je m'appelle ?", keep_history=True) | |
| client.reset() # effacer l'historique | |
| ``` | |
| --- | |
| ## Authentification | |
| Si `CORTEX_API_KEY` est défini, chaque requête doit porter l'en-tête : | |
| ```text | |
| Authorization: Bearer <votre-cle> | |
| ``` | |
| Sans en-tête valide, le serveur répond `401 invalid API key`. | |
| Sans `CORTEX_API_KEY`, aucune authentification n'est demandée. Ne l'exposez | |
| jamais sur Internet dans cet état. | |
| --- | |
| ## Codes d'erreur | |
| | Code | Cause | | |
| |---|---| | |
| | `400` | `messages` vide, ou `stream=true` demandé | | |
| | `401` | Clé absente ou incorrecte | | |
| | `422` | Corps de requête mal formé | | |
| > **`stream=true` n'est pas encore pris en charge.** Le serveur refuse | |
| > explicitement la demande plutôt que de renvoyer une réponse trompeuse. | |
| --- | |
| ## Ajouter un outil | |
| ```python | |
| from cortex_ai.tools import tool, ToolRegistry | |
| from cortex_ai.engine import CortexAgent | |
| from cortex_ai.api import create_app | |
| from cortex_ai.adapters import MockAdapter | |
| @tool(description="Renvoie la longueur d'un texte.") | |
| def longueur(texte: str) -> str: | |
| return str(len(texte)) | |
| agent = CortexAgent(MockAdapter(), ToolRegistry([longueur])) | |
| app = create_app(agent.adapter) | |
| ``` | |
| Le schéma JSON est déduit des annotations de type. Le modèle reçoit | |
| automatiquement la description de l'outil. | |
| --- | |
| ## Déploiement | |
| En production : | |
| ```bash | |
| export CORTEX_API_KEY="une-vraie-cle-longue" | |
| export CORTEX_PORT=8000 | |
| export CORTEX_ADAPTER=hf | |
| python -m uvicorn cortex_ai.api.server:create_app --factory --host 0.0.0.0 --port 8000 | |
| ``` | |
| Placez le serveur derrière un reverse proxy avec TLS. Vérifiez le matériel | |
| disponible : le checkpoint complet demande plusieurs accélérateurs. | |