Instructions to use Blackroot/Ammeg-26B-GGUF-Q6 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Blackroot/Ammeg-26B-GGUF-Q6 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K # Run inference directly in the terminal: llama cli -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K # Run inference directly in the terminal: llama cli -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K # Run inference directly in the terminal: ./llama-cli -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Use Docker
docker model run hf.co/Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
- LM Studio
- Jan
- Ollama
How to use Blackroot/Ammeg-26B-GGUF-Q6 with Ollama:
ollama run hf.co/Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
- Unsloth Desktop
- Pi
How to use Blackroot/Ammeg-26B-GGUF-Q6 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Blackroot/Ammeg-26B-GGUF-Q6:Q6_K" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Blackroot/Ammeg-26B-GGUF-Q6 with Docker Model Runner:
docker model run hf.co/Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
- Lemonade
How to use Blackroot/Ammeg-26B-GGUF-Q6 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Run and chat with the model
lemonade run user.Ammeg-26B-GGUF-Q6-Q6_K
List all available models
lemonade list
- Hermes Agent
How to use Blackroot/Ammeg-26B-GGUF-Q6 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Blackroot/Ammeg-26B-GGUF-Q6 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Blackroot/Ammeg-26B-GGUF-Q6:Q6_K
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Blackroot/Ammeg-26B-GGUF-Q6:Q6_K" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Disclaimer: The specific method of tuning has likely removed the entirety of the models safeguards, including its tendency to soft-refuse and redirect. This also means that safeguards that you'd typically want or expect in a model are probably not present, this is more of a hammer method than a scalpel.
This model is an experimental finetune of gemma-26b-4A ~ it follows the exact same instruct prompting methods as the original model did.
Two parts: A new kind of qualtiy-preserving abliteration (loosely based on heretic) followed by retraining the abliterated model using an evolutionary strategy loosely based off of https://arxiv.org/abs/2511.16652
I've talked about the abliteration in my prior model setup, so I'll discuss only the finetuning method here:
Data, briefly
We start with a baseline sample of a variety of (primarily books), chunk them, and then have the abliterated model caption the stories as a prompt. -> "Generate a story with a protagnoist named Alice..."
Tuning Method, also briefly
The model was abliterated and trained in full (BF16) precision. Low Rank was empoyed on a per-sample basis (Rank 1 LORA per sample into a full-precision buffer -- with stochastic rounding) this means each update is highly-approximated, but the noisy landscape is represented in full-precision buffer, so eventually we get noise cancellation in the buffer and it becomes a full-rank tuning method in a gaussian landscape.
We use a non-differentiable objective combined with teacher forcing. The specific non-differentiable objective is overly complex to describe but the primary part exploits the zipf structure of language as a proxy for long-term dependencies in stories.
Hardware and Software notes
This training was all done purely on some very powerful CPUs I have in my garage running at about 350W and took approximately 280 hours. The cost of power to me comes out to about ~20-25$ USD, the hardware cost is about ~20K USD at the time of this model upload. The main downside of this method is time, it's quite slow. All of this was done in a custom inference engine I'm writing in mojo, including the abliteration and ES training.
- Downloads last month
- 47
6-bit