Instructions to use Denisijcu/Vertex-Core-Gemma with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Denisijcu/Vertex-Core-Gemma with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Denisijcu/Vertex-Core-Gemma")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Denisijcu/Vertex-Core-Gemma", device_map="auto") - PEFT
How to use Denisijcu/Vertex-Core-Gemma with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Denisijcu/Vertex-Core-Gemma with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Denisijcu/Vertex-Core-Gemma:Q4_K_M # Run inference directly in the terminal: llama cli -hf Denisijcu/Vertex-Core-Gemma:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Denisijcu/Vertex-Core-Gemma:Q4_K_M # Run inference directly in the terminal: llama cli -hf Denisijcu/Vertex-Core-Gemma:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Denisijcu/Vertex-Core-Gemma:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Denisijcu/Vertex-Core-Gemma:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Denisijcu/Vertex-Core-Gemma:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Denisijcu/Vertex-Core-Gemma:Q4_K_M
Use Docker
docker model run hf.co/Denisijcu/Vertex-Core-Gemma:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Denisijcu/Vertex-Core-Gemma with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Denisijcu/Vertex-Core-Gemma" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Denisijcu/Vertex-Core-Gemma", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Denisijcu/Vertex-Core-Gemma:Q4_K_M
- SGLang
How to use Denisijcu/Vertex-Core-Gemma with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Denisijcu/Vertex-Core-Gemma" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Denisijcu/Vertex-Core-Gemma", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Denisijcu/Vertex-Core-Gemma" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Denisijcu/Vertex-Core-Gemma", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use Denisijcu/Vertex-Core-Gemma with Ollama:
ollama run hf.co/Denisijcu/Vertex-Core-Gemma:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Denisijcu/Vertex-Core-Gemma with Docker Model Runner:
docker model run hf.co/Denisijcu/Vertex-Core-Gemma:Q4_K_M
- Lemonade
How to use Denisijcu/Vertex-Core-Gemma with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Denisijcu/Vertex-Core-Gemma:Q4_K_M
Run and chat with the model
lemonade run user.Vertex-Core-Gemma-Q4_K_M
List all available models
lemonade list
- Atomic Chat
- Vertex Core
- Intended Use
- Vertex Core Agent Architecture
- Fine-Tuning
- Loading the LoRA Adapter
- Merged Model
- GGUF
- Local Deployment
- Hardware
- Recommended Agent Architecture
- Recommended Engineering Loop
- Example Tool Interface
- Software Engineering Capabilities
- Cybersecurity Use
- Safety and Security
- Limitations
- Known Evaluation Status
- Model Formats
- Reproducibility
- Prompting
- Base Model
- Acknowledgements
- Developer
- Citation
- Project
- License
- Status
Vertex Core
Vertex Core is a specialized AI software-engineering model developed by Vertex Coders LLC.
It is based on Google's Gemma 2 9B IT and fine-tuned using LoRA/QLoRA techniques to improve its behavior for software-engineering workflows, code generation, code modification, debugging, security-oriented code analysis, and tool-oriented development tasks.
Vertex Core is designed to serve as the coding and reasoning core of agentic software-engineering systems.
Status: Experimental / Research Release
Vertex Core is under active development.
Model Overview
| Property | Value |
|---|---|
| Model | Vertex Core |
| Base model | google/gemma-2-9b-it |
| Architecture | Gemma 2 9B |
| Fine-tuning | LoRA / QLoRA |
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0.05 |
| Trainable parameters | ~54 million |
| Target modules | q_proj, k_proj, v_proj, o_proj, up_proj, down_proj, gate_proj |
| Primary use | Software engineering |
| Secondary use | Code and security analysis |
| Primary languages | English, Spanish |
| Developer | Vertex Coders LLC |
Intended Use
Vertex Core is intended for applications involving:
- Software engineering assistance
- Code generation
- Code modification
- Debugging
- Repository analysis
- Code review
- Security-oriented code analysis
- Vulnerability remediation
- Automated development workflows
- Tool-using AI agents
- Local and private AI development environments
- Software-engineering automation
The model is particularly intended to operate as part of an agentic system where the surrounding application provides access to repositories, files, development tools, execution environments, and verification mechanisms.
Vertex Core Agent Architecture
Vertex Core is designed to operate inside an engineering-oriented agent loop such as:
INSPECT
↓
UNDERSTAND
↓
ACT
↓
VERIFY
↓
TEST
↓
CONCLUDE
The model itself does not automatically have access to:
- Filesystems
- Shells
- Networks
- Repositories
- External APIs
- MCP servers
- Development tools
Those capabilities must be provided by the surrounding application or orchestration layer.
The purpose of this architecture is to separate:
Model Reasoning
+
Tool Execution
+
Verification
rather than treating generated text as proof that an operation was successfully performed.
Fine-Tuning
Vertex Core was developed using parameter-efficient fine-tuning techniques based on LoRA / QLoRA.
LoRA Configuration
LoRA rank: 16
LoRA alpha: 32
LoRA dropout: 0.05
Target Modules
q_proj
k_proj
v_proj
o_proj
up_proj
down_proj
gate_proj
The resulting adapter contains approximately:
54 million trainable LoRA parameters
The base model remains:
google/gemma-2-9b-it
Loading the LoRA Adapter
The adapter can be loaded on top of the original Gemma 2 9B IT model using transformers and peft.
Example:
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
BASE_MODEL = "google/gemma-2-9b-it"
LORA_PATH = "PATH_TO_VERTEX_CORE_LORA"
tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL)
base_model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL
)
model = PeftModel.from_pretrained(
base_model,
LORA_PATH
)
For production or local inference, use the loading configuration appropriate for the available hardware.
Merged Model
A merged version of Vertex Core can also be created by merging the LoRA adapter with the original Gemma 2 9B IT base model.
A merged Transformers model is useful when a standalone model directory is preferred over loading the base model and adapter separately.
GGUF
GGUF versions of Vertex Core are intended for local inference environments compatible with the GGUF / llama.cpp ecosystem.
This includes applications such as:
- LM Studio
- llama.cpp-compatible runtimes
- Other local inference applications supporting GGUF
Quantized variants can significantly reduce memory requirements compared with full-precision model weights.
Local Deployment
Vertex Core has been tested in a local development environment using LM Studio.
LM Studio can expose an OpenAI-compatible local inference API.
Example local endpoint:
http://127.0.0.1:1234/v1/chat/completions
This endpoint is intended for local development and is not a public Vertex Core service.
A compatible OpenAI-style client can be used to communicate with the local model server.
Hardware
Vertex Core has been tested in a local development environment using:
GPU:
NVIDIA GTX 1660 Ti
VRAM:
6 GB
Runtime:
CUDA-enabled PyTorch
Local inference:
LM Studio
For GPUs with limited VRAM, quantized GGUF variants are recommended.
Actual performance and memory requirements depend on:
- Quantization level
- Context length
- Batch size
- Runtime
- GPU architecture
- CPU/RAM configuration
Recommended Agent Architecture
For autonomous or semi-autonomous software-engineering applications, Vertex Core should be combined with an execution and verification layer.
A recommended architecture is:
┌──────────────────────┐
│ User Issue │
└──────────┬───────────┘
↓
┌──────────────────────┐
│ Vertex Core │
│ Reasoning / Coding │
└──────────┬───────────┘
↓
┌──────────────────────┐
│ Agent Orchestrator │
└──────────┬───────────┘
↓
┌────────────────────────────────┐
│ Tools │
│ │
│ read_file │
│ write_file │
│ replace_in_file │
│ list_directory │
│ run_command │
│ run_tests │
│ security_scan │
└────────────────┬───────────────┘
↓
┌──────────────────────┐
│ Verification │
└──────────┬───────────┘
↓
┌──────────────────────┐
│ Conclusion │
└──────────────────────┘
The orchestration layer should independently verify important operations.
Recommended Engineering Loop
A practical software-engineering workflow is:
User Issue
↓
Inspect Repository
↓
Read Relevant Files
↓
Understand Problem
↓
Determine Required Change
↓
Modify Code
↓
Read Modified Code
↓
Run Tests
↓
Verify Result
↓
Report Outcome
The verification stage is especially important for agentic systems.
A model response claiming that a file was modified should not be treated as proof that the file was actually changed.
Example Tool Interface
A surrounding agent system may provide tools such as:
read_file
write_file
replace_in_file
list_directory
run_command
run_tests
security_scan
For example, the orchestration layer may expose a repository workspace:
/sandbox
and allow Vertex Core to reason about the contents of that workspace through controlled tools.
The exact tool implementation is application-specific.
Software Engineering Capabilities
Vertex Core is designed to assist with tasks such as:
Code Generation
Creating functions, classes, modules, utilities, tests, and other software components.
Code Modification
Changing existing implementations according to explicit requirements.
Debugging
Analyzing errors, stack traces, implementation details, and behavioral problems.
Repository Analysis
Inspecting project structure and identifying relevant files before making changes.
Security-Oriented Development
Analyzing code for common security problems and proposing remediation.
Agentic Development
Working as the reasoning component of a larger system capable of reading files, executing commands, running tests, and verifying changes.
Cybersecurity Use
Vertex Core can be integrated into defensive cybersecurity workflows involving:
- Secure code review
- Vulnerability analysis
- Security-oriented debugging
- Static-analysis interpretation
- Remediation assistance
- Defensive development workflows
Security capabilities should only be used on systems, networks, repositories, and applications that the user is authorized to analyze or modify.
Vertex Core should not be considered a replacement for:
- Professional penetration testing
- Security audits
- Code review
- Vulnerability management
- Operational security controls
- Human security expertise
Safety and Security
Vertex Core is a language model and can produce incorrect, incomplete, or unsafe outputs.
Applications integrating the model should implement appropriate controls around:
- Tool permissions
- Filesystem access
- Command execution
- Network access
- Secrets
- Credentials
- Repository access
- Production systems
For autonomous systems, potentially destructive operations should be restricted or require explicit authorization.
Limitations
Vertex Core is an experimental fine-tuned language model.
It can:
- Generate incorrect code
- Misunderstand requirements
- Hallucinate APIs
- Hallucinate files
- Produce invalid commands
- Misinterpret tool output
- Produce incomplete modifications
- Fail to recognize an issue
- Suggest an incorrect security remediation
Generated code must therefore be tested before deployment.
Tool execution results should be independently verified.
A successful shell command does not necessarily mean that the requested semantic operation was successfully completed.
For this reason, agent implementations should verify filesystem state and test relevant behavior after modifications.
Security findings should also be independently validated.
Known Evaluation Status
Vertex Core has been evaluated in local software-engineering workflows involving:
- Code generation
- Code modification
- SQL injection remediation
- Repository inspection
- Tool-oriented agent workflows
- Local inference through LM Studio
Testing indicates that the model can perform straightforward code transformations and security-oriented code modifications.
However, the current release should not be considered a fully reliable autonomous software-engineering agent.
The surrounding orchestration and verification layer remains an important part of the system.
Model Formats
Vertex Core can be distributed in several forms.
LoRA Adapter
Recommended for users who want to reproduce the fine-tuning setup on top of the original Gemma model.
Vertex Core LoRA
+
Gemma 2 9B IT
↓
Vertex Core
Merged Transformers Model
Recommended for applications that prefer a standalone Transformers model.
GGUF
Recommended for local inference applications supporting GGUF.
Quantized versions are particularly useful for consumer GPUs with limited VRAM.
Reproducibility
To reproduce the adapter-based model, users should obtain:
Base Model:
google/gemma-2-9b-it
Vertex Core:
LoRA Adapter
Runtime:
Transformers + PEFT
The exact inference behavior may vary depending on:
- Transformers version
- PEFT version
- Quantization configuration
- Hardware
- Sampling parameters
- Prompt formatting
- Chat template
- Runtime implementation
Prompting
Vertex Core is intended to work best when the surrounding application clearly defines:
- The task
- The available tools
- The workspace
- The expected behavior
- The verification requirements
For agentic applications, the model should be instructed to distinguish between:
Intent
↓
Observation
↓
Action
↓
Tool Result
↓
Verification
↓
Conclusion
This helps reduce the risk of treating an imagined action as a completed operation.
Base Model
Vertex Core is based on:
Google Gemma 2 9B IT
The original base model is developed by Google and remains subject to its respective license, terms of use, and policies.
Users of Vertex Core should review the applicable Gemma licensing and usage requirements.
Acknowledgements
Vertex Coders acknowledges the work of the Google Gemma team and the open-source machine-learning ecosystem that makes parameter-efficient fine-tuning and local model deployment possible.
Vertex Core builds upon:
- Google Gemma
- Hugging Face Transformers
- Hugging Face PEFT
- LoRA / QLoRA
- GGUF-compatible inference technologies
- LM Studio
Developer
Vertex Coders LLC
Vertex Core is part of the Vertex Coders AI engineering ecosystem.
The project is developed as part of Vertex Coders' work on agentic software engineering, AI-powered development systems, cybersecurity tooling, and local AI infrastructure.
Citation
If you use Vertex Core in research, demonstrations, evaluations, or derivative projects, please reference:
Vertex Coders LLC.
Vertex Core — AI Software Engineering Agent.
Vertex Coders, 2026.
A formal academic citation will be provided in a future release.
Project
Source code and engineering infrastructure:
Vertex Coders — Vertex Core
The model and the surrounding agent infrastructure are maintained as separate components so that the model can be distributed independently from the execution and orchestration layer.
License
This repository contains the Vertex Core model adapter and associated documentation.
The adapter is released under the license specified in this model repository.
The underlying Gemma 2 9B IT model remains subject to Google's applicable Gemma terms and license.
Users are responsible for complying with the licenses and terms applicable to both the base model and any additional components used with Vertex Core.
Status
Experimental / Research Release
Vertex Core is under active development.
Future releases may improve:
- Software-engineering reliability
- Tool-use behavior
- Repository reasoning
- Code modification accuracy
- Verification behavior
- Security analysis
- Multilingual performance
- Local inference efficiency
- Agent orchestration
Una corrección importante antes de subirlo
Hay una cosa que no quiero que hagamos a ciegas: el license: apache-2.0 que pusimos arriba.
Si el adapter de Vertex Core realmente lo quieres distribuir bajo Apache-2.0, perfecto. Pero eso no convierte a Gemma 2 en Apache-2.0. La sección que dejé al final separa explícitamente ambas cosas.
Para la publicación, yo usaría este README como la Model Card oficial de Vertex Core y subiría el adapter_model.safetensors junto con adapter_config.json. Los GGUF los podemos publicar después como archivos/release separados.
- Downloads last month
- 8
4-bit