Text Generation
Transformers
Safetensors
French
English
Chinese
deepseek_v4
cortex
code-generation
web-development
software-engineering
Mixture of Experts
8-bit precision
fp8
Instructions to use Frankenstein-Labs/cortex.6.sol with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Frankenstein-Labs/cortex.6.sol with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Frankenstein-Labs/cortex.6.sol")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Frankenstein-Labs/cortex.6.sol") model = AutoModelForCausalLM.from_pretrained("Frankenstein-Labs/cortex.6.sol", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Frankenstein-Labs/cortex.6.sol with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Frankenstein-Labs/cortex.6.sol" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Frankenstein-Labs/cortex.6.sol
- SGLang
How to use Frankenstein-Labs/cortex.6.sol with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/cortex.6.sol" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Frankenstein-Labs/cortex.6.sol" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frankenstein-Labs/cortex.6.sol", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Frankenstein-Labs/cortex.6.sol with Docker Model Runner:
docker model run hf.co/Frankenstein-Labs/cortex.6.sol
Download cortex_ai/engine/dsml.py from Frankenstein-Labs/cortex.6.sol: direct link, hf CLI and curl.
- Browser
- Download file 3.01 kB
-
https://huggingface.co/Frankenstein-Labs/cortex.6.sol/resolve/main/cortex_ai/engine/dsml.py
- Command line
-
hf download hf://Frankenstein-Labs/cortex.6.sol/cortex_ai/engine/dsml.py
-
curl -L -o dsml.py https://huggingface.co/Frankenstein-Labs/cortex.6.sol/resolve/main/cortex_ai/engine/dsml.py
3.01 kB
| """Builders for well-formed DeepSeek-V4 completion text. | |
| The encoding module defines a strict format: ``parse_message_from_completion_text`` | |
| validates every token and raises on the slightest deviation. Anything that | |
| produces completions outside the model -- a mock adapter, a replay file, a | |
| conformance test -- should build them with these helpers rather than by hand. | |
| The special tokens are imported from ``encoding_dsv4`` rather than redefined | |
| here. They contain non-ASCII codepoints that do not survive copy-paste, and a | |
| single wrong byte makes every parse fail. | |
| """ | |
| from __future__ import annotations | |
| import json | |
| from typing import Any | |
| def _tokens() -> dict[str, str]: | |
| try: | |
| import encoding_dsv4 as enc | |
| except ImportError as exc: # pragma: no cover - environment dependent | |
| raise RuntimeError( | |
| "The encoding module is required. Add the repository's encoding/ folder " | |
| "to PYTHONPATH, e.g. PYTHONPATH=/workspace/project/encoding" | |
| ) from exc | |
| return { | |
| "dsml": enc.dsml_token, | |
| "eos": enc.eos_token, | |
| "thinking_start": enc.thinking_start_token, | |
| "thinking_end": enc.thinking_end_token, | |
| "tool_calls_block": enc.tool_calls_block_name, | |
| } | |
| def _parameter(name: str, value: Any) -> str: | |
| dsml = _tokens()["dsml"] | |
| is_str = isinstance(value, str) | |
| rendered = value if is_str else json.dumps(value, ensure_ascii=False) | |
| flag = "true" if is_str else "false" | |
| return f'<{dsml}parameter name="{name}" string="{flag}">{rendered}</{dsml}parameter>' | |
| def build_tool_call(name: str, arguments: dict[str, Any], reasoning: str = "") -> str: | |
| """Render one assistant turn that calls a single tool. | |
| Reproduces the exact byte sequence the parser expects: a leading blank line, | |
| one parameter per line, the closing invoke tag, and the EOS token. | |
| In thinking mode the completion opens with the reasoning content and closes | |
| it with the end tag. The opening `` thinking`` token belongs to the *prompt*, | |
| so it must not appear here; the parser rejects it inside content. | |
| """ | |
| t = _tokens() | |
| lines = [f'<{t["dsml"]}invoke name="{name}">'] | |
| for key, value in arguments.items(): | |
| lines.append(_parameter(key, value)) | |
| body = "\n".join(lines) + f"\n</{t['dsml']}invoke>" | |
| call = ( | |
| f"\n\n<{t['dsml']}{t['tool_calls_block']}>\n{body}\n" | |
| f"</{t['dsml']}{t['tool_calls_block']}>" | |
| f"{t['eos']}" | |
| ) | |
| if reasoning: | |
| return f"{reasoning}{t['thinking_end']}{call}" | |
| return call | |
| def build_answer(content: str, reasoning: str = "") -> str: | |
| """Render one assistant turn that answers without calling a tool.""" | |
| t = _tokens() | |
| if reasoning: | |
| return f"{reasoning}{t['thinking_end']}{content}{t['eos']}" | |
| return f"{content}{t['eos']}" | |
| def build_tool_result(content: str) -> str: | |
| """Render a tool result body, matching the encoder's tool_output_template.""" | |
| return f"<tool_result>{content}</tool_result>" | |