Text Generation
Transformers
Safetensors
French
English
Chinese
deepseek_v4
cortex
code-generation
web-development
software-engineering
Mixture of Experts
8-bit precision
fp8
Instructions to use Frankenstein-Labs/Cortex-ai with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Frankenstein-Labs/Cortex-ai with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Frankenstein-Labs/Cortex-ai")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Frankenstein-Labs/Cortex-ai") model = AutoModelForCausalLM.from_pretrained("Frankenstein-Labs/Cortex-ai", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Frankenstein-Labs/Cortex-ai with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Frankenstein-Labs/Cortex-ai" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/Cortex-ai", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Frankenstein-Labs/Cortex-ai
- SGLang
How to use Frankenstein-Labs/Cortex-ai with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/Cortex-ai" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/Cortex-ai", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/Cortex-ai" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/Cortex-ai", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Frankenstein-Labs/Cortex-ai with Docker Model Runner:
docker model run hf.co/Frankenstein-Labs/Cortex-ai
Download cortex_ai/api/server.py from Frankenstein-Labs/Cortex-ai: direct link, hf CLI and curl.
- Browser
- Download file 6.68 kB
-
https://huggingface.co/Frankenstein-Labs/Cortex-ai/resolve/main/cortex_ai/api/server.py
- Command line
-
hf download hf://Frankenstein-Labs/Cortex-ai/cortex_ai/api/server.py
-
curl -L -o server.py https://huggingface.co/Frankenstein-Labs/Cortex-ai/resolve/main/cortex_ai/api/server.py
6.68 kB
| """OpenAI-compatible HTTP API for CORTEX AI. | |
| Any client that speaks the OpenAI chat-completions protocol -- the official | |
| Python SDK, LangChain, LlamaIndex, a curl script -- can talk to CORTEX AI by | |
| changing only the base URL. | |
| Endpoints: | |
| GET /health | |
| GET /v1/models | |
| POST /v1/chat/completions | |
| """ | |
| from __future__ import annotations | |
| import time | |
| import uuid | |
| from typing import Any | |
| from fastapi import FastAPI, Header, HTTPException | |
| from pydantic import BaseModel, Field | |
| from ..adapters.base import ModelAdapter | |
| from ..config import CortexConfig | |
| from ..engine.agent import CortexAgent | |
| from ..identity import identity_language, identity_response, is_identity_question | |
| from ..tools.registry import registry_from_names | |
| class ChatMessage(BaseModel): | |
| role: str | |
| content: str = "" | |
| name: str | None = None | |
| class ChatCompletionRequest(BaseModel): | |
| model: str | None = None | |
| messages: list[ChatMessage] | |
| temperature: float | None = None | |
| max_tokens: int | None = None | |
| stream: bool = False | |
| tools: list[dict[str, Any]] | None = None | |
| class Usage(BaseModel): | |
| prompt_tokens: int = 0 | |
| completion_tokens: int = 0 | |
| total_tokens: int = 0 | |
| class ChatCompletionChoice(BaseModel): | |
| index: int = 0 | |
| message: ChatMessage | |
| finish_reason: str = "stop" | |
| class ChatCompletionResponse(BaseModel): | |
| id: str | |
| object: str = "chat.completion" | |
| created: int | |
| model: str | |
| choices: list[ChatCompletionChoice] | |
| usage: Usage | |
| # CORTEX extension: the reasoning trace, when thinking mode is on. | |
| reasoning_content: str = "" | |
| tool_calls: list[dict[str, Any]] = Field(default_factory=list) | |
| def create_app(adapter: ModelAdapter, config: CortexConfig | None = None) -> FastAPI: | |
| """Build the FastAPI application around a model adapter.""" | |
| cfg = config or CortexConfig() | |
| tools = registry_from_names(cfg.enabled_tools) | |
| agent = CortexAgent( | |
| adapter, | |
| tools, | |
| cfg.engine, | |
| system_prompt=cfg.system_prompt, | |
| ) | |
| app = FastAPI( | |
| title="CORTEX AI API", | |
| version="1.0.0", | |
| description="API compatible OpenAI pour CORTEX AI, un projet de Frankenstein-Labs.", | |
| ) | |
| app.state.cortex_config = cfg | |
| app.state.cortex_agent = agent | |
| app.state.identity_interception = True | |
| def _check_auth(authorization: str | None) -> None: | |
| if not cfg.server.requires_auth: | |
| return | |
| expected = f"Bearer {cfg.server.api_key}" | |
| if authorization != expected: | |
| raise HTTPException(status_code=401, detail="invalid API key") | |
| def health() -> dict[str, Any]: | |
| return { | |
| "status": "ok", | |
| "model": cfg.model_id, | |
| "tools": tools.names(), | |
| "thinking_mode": cfg.engine.thinking_mode, | |
| "identity_interception": "deterministic", | |
| } | |
| def list_models(authorization: str | None = Header(default=None)) -> dict[str, Any]: | |
| _check_auth(authorization) | |
| return { | |
| "object": "list", | |
| "data": [ | |
| { | |
| "id": cfg.model_id, | |
| "object": "model", | |
| "created": int(time.time()), | |
| "owned_by": "Frankenstein-Labs", | |
| } | |
| ], | |
| } | |
| def chat_completions( | |
| request: ChatCompletionRequest, | |
| authorization: str | None = Header(default=None), | |
| ) -> ChatCompletionResponse: | |
| _check_auth(authorization) | |
| if request.stream: | |
| raise HTTPException( | |
| status_code=400, | |
| detail="stream=true is not supported yet; use stream=false", | |
| ) | |
| if not request.messages: | |
| raise HTTPException(status_code=400, detail="messages must not be empty") | |
| # Deterministic identity boundary: answer before system prompts, tools, | |
| # or model inference can alter the canonical creator attribution. | |
| last_user_message = next( | |
| (m.content for m in reversed(request.messages) if m.role == "user"), | |
| None, | |
| ) | |
| if last_user_message is not None and is_identity_question(last_user_message): | |
| content = identity_response(identity_language(last_user_message)) | |
| prompt_tokens = adapter.count_tokens(last_user_message) | |
| completion_tokens = adapter.count_tokens(content) | |
| return ChatCompletionResponse( | |
| id=f"chatcmpl-{uuid.uuid4().hex[:24]}", | |
| created=int(time.time()), | |
| model=request.model or cfg.model_id, | |
| choices=[ | |
| ChatCompletionChoice( | |
| index=0, | |
| message=ChatMessage(role="assistant", content=content), | |
| finish_reason="stop", | |
| ) | |
| ], | |
| usage=Usage( | |
| prompt_tokens=prompt_tokens, | |
| completion_tokens=completion_tokens, | |
| total_tokens=prompt_tokens + completion_tokens, | |
| ), | |
| ) | |
| system_override: list[dict[str, Any]] = [] | |
| turns: list[dict[str, Any]] = [] | |
| for m in request.messages: | |
| if m.role == "system": | |
| system_override.append({"role": "system", "content": m.content}) | |
| else: | |
| turns.append({"role": m.role, "content": m.content}) | |
| if system_override: | |
| agent.system_prompt = system_override[-1]["content"] | |
| result = agent.run(turns) | |
| prompt_tokens = sum(adapter.count_tokens(m["content"]) for m in turns) | |
| completion_tokens = adapter.count_tokens(result.content) | |
| return ChatCompletionResponse( | |
| id=f"chatcmpl-{uuid.uuid4().hex[:24]}", | |
| created=int(time.time()), | |
| model=request.model or cfg.model_id, | |
| choices=[ | |
| ChatCompletionChoice( | |
| index=0, | |
| message=ChatMessage(role="assistant", content=result.content), | |
| finish_reason="stop", | |
| ) | |
| ], | |
| usage=Usage( | |
| prompt_tokens=prompt_tokens, | |
| completion_tokens=completion_tokens, | |
| total_tokens=prompt_tokens + completion_tokens, | |
| ), | |
| reasoning_content=result.reasoning, | |
| tool_calls=[ | |
| {"name": c.name, "arguments": c.arguments, "result": c.result, "ok": c.ok} | |
| for c in result.tool_calls | |
| ], | |
| ) | |
| return app | |