GLM 5.2 orchestrates, DeepSeek V4 Flash executes
How to design multi-model systems where one open LLM plans, another executes, and the context they hand each other stays efficient.
15 articles
How to design multi-model systems where one open LLM plans, another executes, and the context they hand each other stays efficient.
GLM 5.2, the 744B MIT open model, now runs on Helmcode's European infrastructure with zero logs and a flat rate of 150 euros a month.
The calm version of the AI Act: which obligations really apply to you from 2 August, which ones the Omnibus moved, and what to have ready.
Upload your documentation once or connect a Git repository, then ask it from the chat. Every answer cites the document it came from.
FP8, NVFP4, H200, B200: what each NVIDIA generation accelerates in hardware, how much VRAM your model needs, and which GPU to buy. With real production data.
The difference between an image that screams 'AI' and a professional one isn't the model, it's the prompt. The complete guide to generating great images in Helmcode.
New hire at Helmcode: communication, marketing and community. Who I am, what I'm here to do, and what we're planning for the NaN community.
How we moved Qwen3.6-35B to NVFP4 on a single RTX PRO 6000, why it kept crashing, and the one-line fix that turned out to be a cuDNN bug, not a vLLM bug.
117 billion tokens, 3.68 million requests, 21 countries, and 99.98% uptime. NaN is a community of builders with its own inference infrastructure and a private platform to deploy apps and agents.
In this post we'll learn what parameters and quantization are, so we can figure out how much space AI models take up.
In this post I'll walk you through how the community's inference servers are set up: the hardware we use, the stack we run, and the models we serve.
I've spent several hours over several days documenting and optimizing my entire local environment so I can "mechanize" the work I do every day managing infrastructure for multiple startups.
This post isn't meant to be a guide on how to use Clawd, but rather a look at how we're rolling it out at Helmcode to have an AI Agent that helps us with our day-to-day work managing the Cloud infrastructure of multiple startups.
Kubernetes is one of the most widely used infrastructure tools among companies, and it has become the standard for running containerized applications at scale all over the world.
Before we start, a bit of context. The infrastructure is hosted on AWS and the architecture was based on Serverless services:
// cookies
We only use strictly necessary cookies to run the site. No analytics, no advertising, ever — see our Cookie Policy.
// preferences