Instructions to use Brunobkr/OFFELLIA_Quantis with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Brunobkr/OFFELLIA_Quantis with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M # Run inference directly in the terminal: llama cli -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M # Run inference directly in the terminal: llama cli -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M
Use Docker
docker model run hf.co/Brunobkr/OFFELLIA_Quantis:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Brunobkr/OFFELLIA_Quantis with Ollama:
ollama run hf.co/Brunobkr/OFFELLIA_Quantis:Q4_K_M
- Unsloth Desktop
- Pi
How to use Brunobkr/OFFELLIA_Quantis with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Brunobkr/OFFELLIA_Quantis:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Brunobkr/OFFELLIA_Quantis with Docker Model Runner:
docker model run hf.co/Brunobkr/OFFELLIA_Quantis:Q4_K_M
- Lemonade
How to use Brunobkr/OFFELLIA_Quantis with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Brunobkr/OFFELLIA_Quantis:Q4_K_M
Run and chat with the model
lemonade run user.OFFELLIA_Quantis-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Brunobkr/OFFELLIA_Quantis with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Brunobkr/OFFELLIA_Quantis:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Brunobkr/OFFELLIA_Quantis with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Brunobkr/OFFELLIA_Quantis:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Brunobkr/OFFELLIA_Quantis:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
ΩFFFΣLLIa • llama.cpp • OFFFELLIA_PURE
██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗
██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗
██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║
██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║
╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║-PURE
╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝
High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem in Pure C/C++
📖 Visão Geral
ΩFFFΣLLIa • llama.cpp • AlgMor24 é um fork avançado, destravado e de alta performance do ecossistema llama.cpp. Este projeto integra inferência local de última geração em C/C++ com um motor agêntico autônomo multi-turn, suporte nativo a FIM (Fill-in-the-Middle) para geração e preenchimento de código, Speculative Decoding otimizado para programação, integração de ferramentas MCP (Model Context Protocol) e uma interface Web moderna em SvelteKit/Vite com a identidade visual Cyberpunk Neon Fire.
✨
git clone https://github.com/brunoconta1980-tech/llama_OFFFELLIA_1984
cd llama_OFFFELLIA_1984
cmake -B build
-DGGML_VULKAN=ON
-DLLAMA_BUILD_WEBUI=ON
-DLLAMA_SERVER_TOOLS=ON
cmake --build build -j
or
cmake -S . -B build-vulkan
-DGGML_VULKAN=ON
-DLLAMA_BUILD_WEBUI=ON
-DLLAMA_SERVER_TOOLS=ON
cmake --build build-vulkan -j
Comando sugerido:
"/home/userk21/llama_OFFFELLIA_1984/build/bin/llama-server"
-m "/home/userk21/Área de trabalho/userk21/LLMS/ΩFFFΣLLIα_IQ4_NL_gemma-4-26B-A4B-it.gguf"
-ngl 99 --n-cpu-moe 99
-c 50000
-ctk q8_0
-ctv q8_0
-t 4
-tb 4
-b 2048
-ub 1024
-fa on
--cpu-strict 1
--parallel 1
--agent
--tools all
--reasoning auto
--kv-unified
--load-mode mmap
--cors-origins "*"
--webui-mcp-proxy
--threads-http -1
--port 5173
--host 127.0.0.1
📜 Licença
Distribuído sob a licença MIT. Veja o arquivo LICENSE para mais detalhes. Gracias https://github.com/charlie12345/ROCmFPX
- Downloads last month
- 1,686