Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
💼
Hiring
227.7
TFLOPS
Cahlen Humphreys
PRO
cahlen
4
85
255
Follow
jmfi's profile picture
webxos's profile picture
9tryV's profile picture
75 followers
·
332 following
https://bigcompute.science
cahlen
AI & ML interests
☠️💻
Recent Activity
posted
an
update
2 days ago
I published the serving setup and benchmark results I’ve been using for GLM-5.3-Flash UD-IQ1_S on a single NVIDIA DGX Spark. The biggest finding was not throughput — it was reliability. With llama.cpp’s default unrestricted reasoning, a real ~18K-token OpenCode request with 53 tools failed to produce any actionable output in 9/20 runs. Adding: --reasoning-budget 2048 changed that to 20/20 successful tool-call responses. A few other measured results: * 27.5–29 tok/s decode with MTP vs 18.7 without * MTP depth 2 outperformed the default depth 3 * 131K context uses ~90.9 GiB resident memory * 256K context also works on the Spark * CPU MoE offload does not meaningfully free memory on GB10 unified memory * newer llama.cpp builds were slightly faster, but introduced tool-call serialization failures, so the repo pins the stable commit Everything in the README is backed by the benchmark scripts and raw results in the repo. If you're running GLM-5.3-Flash on a DGX Spark for agentic coding, this should give you a solid starting point. https://github.com/cahlen/glm-5.3-flash-GGUF-1bit-dgx-spark https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF
liked
a model
6 days ago
unsloth/Qwen3.8-Flash-Next-GGUF
liked
a model
6 days ago
zai-org/GLM-5.3
View all activity
Organizations
cahlen
's models
29
Sort: Recently updated
cahlen/bigcompute-cuda-kernels
Other
•
Updated
May 31
cahlen/keeloq-neural-distinguishers
Updated
Apr 23
cahlen/ramanujan-machine-cuda
Updated
Apr 14
cahlen/class-numbers-cuda
Updated
Apr 14
cahlen/prime-convergents-cuda
Updated
Apr 14
cahlen/zaremba-transfer-operator-cuda
Updated
Apr 14
cahlen/flint-hills-cuda
Updated
Apr 14
cahlen/zaremba-density-cuda
Updated
Apr 14
cahlen/zaremba-cayley-cuda
Updated
Apr 14
cahlen/kronecker-cuda
Updated
Apr 14
cahlen/ramsey-r55-cuda
Updated
Apr 14
cahlen/lyapunov-spectrum-cuda
Updated
Apr 14
cahlen/minkowski-spectrum-cuda
Updated
Apr 14
cahlen/zaremba-transitivity-cuda
Updated
Apr 14
cahlen/hausdorff-spectrum-cuda
Updated
Apr 14
cahlen/erdos-straus-cuda
Updated
Apr 14
cahlen/Convergent-7B
Text Generation
•
8B
•
Updated
Apr 9
•
15
•
1
cahlen/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-GGUF
35B
•
Updated
Apr 3
•
1.14k
•
2
cahlen/qwen3.5-35b-a3b-compacted-GGUF
23B
•
Updated
Mar 28
•
5.64k
•
2
cahlen/pair3-baseline-7b
Updated
Mar 18
cahlen/cofrgenet-f
Text Generation
•
Updated
Mar 7
cahlen/steerling-8b-combined-lora
Text Generation
•
Updated
Mar 5
•
1
•
3
cahlen/black-metal-art-sdxl-lora
Text-to-Image
•
Updated
Mar 1
•
23
•
•
3
cahlen/oc-punk-flyer-sdxl-lora
Text-to-Image
•
Updated
Mar 1
•
15
•
•
2
cahlen/acid-ansi-lora
Text-to-Image
•
Updated
Feb 26
•
5
•
•
1
cahlen/lingbot-world-base-cam-nf4
Image-to-Video
•
Updated
Feb 3
•
20
cahlen/tinyllama-offline-practical-skills-qa-qlora
Question Answering
•
Updated
Apr 17, 2025
•
2
•
1
cahlen/tinyllama-motorcycle-repair-qa-adapter
Text Generation
•
Updated
Apr 15, 2025
•
2
cahlen/setfit-navigation-instructions
Text Classification
•
0.1B
•
Updated
Mar 31, 2025
•
9