Codendum

Shared, local infrastructure for coding agents, for classrooms and teams. A single NVIDIA GB10 (DGX Spark) runs vLLM with an open-weight model, and users work with OpenCode from their own workstations. Developed by Stefano Noferi.

AIOpen SourcevLLMOpenCodeDGX SparkApache 2.0

What Codendum is

Codendum lets an organization share a single machine that runs inference with open-weight models for generative coding. An NVIDIA GB10 system (DGX Spark or equivalent) runs vLLM with a coding model, and concurrent users run OpenCode on their own workstations. It works the same for a classroom or a company team: prompts and source code stay on the organization’s network.

The repository contains scripts, example configuration and documentation to install the service, secure it, test it, measure it and operate it.

How it works

  • The GB10 only serves inference. Code, version control, compilers, builds and tests stay on users’ workstations or in isolated development environments, whatever the language.
  • Everything runs in Docker. vLLM listens on 127.0.0.1 only; the only way in is an nginx container that accepts HTTPS on a dedicated port from the LAN or VPN, checks a per-user API key, applies per-user limits and forwards only /v1/chat/completions and /v1/models.
  • OpenCode uses an OpenAI-compatible provider, with tool calling enabled explicitly and a context limit that matches the server profile.

The reference model is Qwen3-Coder-30B-A3B-Instruct-FP8; container images and the model are pinned by digest.

Measurements

The Benchmarks page of the documentation reports measurements taken on a GB10 with a simulated class of concurrent OpenCode users, as a reference for that model and client version.

Stated limits

A single machine with no high availability; no per-user token quotas; the shared prefix cache opens a possible timing side channel; code execution permissions depend on the client configuration. The threat model and responsibilities are described on the Security and limits page of the documentation.

Roadmap

A possible next step is governance of the content that goes through the model (personal-data redaction, a prompt-injection firewall, agent loop breaking, a tamper-evident audit log) by integrating Admina as a gateway between the proxy and vLLM. The integration is not implemented yet.

noze’s role

Codendum is created and maintained by Stefano Noferi, founder of noze. Version 0.1.0 was released on 29 September 2026, followed by 0.1.1 on 1 October 2026. The release story is on Stefano Noferi’s blog.

License

Codendum is released under the Apache License 2.0. Container images, model weights and client software are downloaded separately and keep their own licenses.

Need support?Under attack?Service Status
Need support?Under attack?Service Status