
Interesting Finds — 2026-10-01
Six notes: a CAD skill library as a Codex plugin, an open source Blender fork with an agent inside, local inference numbers for mimo v2.6 flash, decision models that answer without generating, a GPU cloud report card, and a real-time pathtracer milestone.
Each is a separate find. Editorial takes are mine where noted.
1. text-to-cad — CAD skills as a Codex plugin
earthtojake/text-to-cad (github.com, MIT, free, local): a library of agent skills for CAD and fabrication, now installable as a Codex plugin for Codex 0.142.0 and newer. Install with codex plugin marketplace add earthtojake/text-to-cad followed by codex plugin add text-to-cad@earthtojake, then restart Codex. Twelve skills cover plain-language to 3D model with STEP as the main output plus STL, 3MF, and GLB; off-the-shelf STEP part sourcing; dimensioned PDF engineering drawings; DXF profiles and cut layouts; URDF, SRDF, and SDF robot and simulation files; SendCutSend pre-upload checks; printability and manufacturability review; OrcaSlicer slicing; and Bambu Lab print dispatch. Kernel is build123d on Open CASCADE. Repo HEAD is v0.7.10 at time of writing.
Meaning: the constraint removed is CAD fluency. The operator describes the part; the skill owns the kernel calls, the export formats, and the vendor handoff checks. Same pattern as the decision models below: structured artifacts out, chat text nowhere in the loop. Nothing here was executed locally, so output correctness and viewer quality remain unverified.
2. Mixar — a Blender fork with an agent inside
Mixar-AI/mixar-app (github.com, ~361 stars): an open source AI 3D editor built as a custom fork of Blender 5.2, shipped as a standalone desktop app for Windows, macOS, and Linux. It preserves Blender tooling, shortcuts, and .blend files and adds native editor spaces plus an in-app chat agent that plans and executes multi-step scene work: modeling from prompts, texture painting, materials, UV fixes, retopology, baking, and export prep. Mixar-original code is GPL-3.0-or-later; the backend AI service is closed source and AI features need an account or a bring-your-own key.
Meaning: the interesting move is keeping the artifact editable. Text-to-3D output that lands as a dead mesh is a render; output that lands inside a Blender session with layers and history is a starting point. Retopology, auto-UV, and baking are the unglamorous steps that decide whether generated geometry survives contact with a real pipeline, so those are the claims to verify first. Doc page bodies failed to load during research, so quality evidence is thin.
3. mimo v2.6 flash on an M5 Ultra — long-context local numbers
Xiaomi’s MiMo-V2.6-Flash-RL (huggingface.co, MIT): a 309B total, 15B active sparse mixture-of-experts model with a 1M token context window. Reported local runs on an M5 Ultra with 256GB unified memory (mxfp4, mlx-vlm) show decode at 65.70 tok/s at 8k context, 59.43 at 32k, 44.49 at 128k, and 37.68 at 200k, with prefill around 8.4 seconds at short context and growing from there. The pattern is gradual decode decline with prefill dominating total latency at long context.
Meaning: the numbers that matter are the slopes, not the headline. A 43 percent decode drop from 8k to 200k with prefill as the larger cost is the shape of memory-bound long-context inference on unified memory. All throughput figures are single-source from timeline posts with no harness version, quantization recipe, or thermal data, so treat them as directional. The corroborating facts are architectural: 15B active parameters, sliding-window heavy attention, and MIT weights that permit the port exist independently of the benchmark.
4. vLLM Decision 2.0 — answers without generation
vLLM Semantic Router team (huggingface.co, six models, Apache 2.0): open decision models that take an input plus structured questions and return a probability per answer in a single forward pass. No text generation. Sizes run Kai-0.6B, Eos-0.8B, Sol-2B, Nox-4B, Lux-9B, and Vega-27B. Published single-question latencies on one GPU range from 4.9 ms for Kai to 71.4 ms for Vega. The announced 64 questions in 63 ms figure is consistent with batched single-pass answering on a small model but the exact model, GPU, and method are unconfirmed.
Meaning: routing and classification priced as infrastructure instead of as chat. A ticket triage or urgency score that costs one forward pass instead of a generated rationale changes what can sit in the request path. Connects to the System One writeup from the September 16 batch: the field is converging on calibrated small models for decisions and large models for everything else. The shared-team benchmark behind the eval tables needs independent reproduction before the rankings mean much.
5. Vultr and ClusterMAX — a GPU cloud report card
SemiAnalysis ClusterMAX (semianalysis.com): hands-on ratings of GPU cloud providers across setup, performance, reliability, security, support, and pricing. Version 3.0 from September 2026 rates 77 of 323 tracked providers with burn-in and fault-injection testing. The documented Vultr findings from the public 2.0 review (Silver tier) include missing cluster tooling at handoff, link flaps with no proactive notification or remediation, and year-old operator versions subject to critical CVEs including NVIDIAscape (CVE-2025-23266, independently verified, fixed upstream). Package-level CVE claims about current Vultr images appear only in timeline posts and are unconfirmed in public text, as is Vultr’s exact 3.0 tier.
Meaning: the bottleneck being measured is operations, not silicon. Bandwidth numbers converge across providers; monitoring, image currency, and incident response do not. A one-year-old GPU operator with a known container escape is the kind of finding that matters more than a percentage point of NCCL bandwidth. Note the standing caveat: one benchmarking firm also sells consulting, so second-source the technical claims where possible.
6. A real-time three.js pathtracer reaches milestone 2
Jeffrey Castellano (@CastellanoWorks): milestone 2 of a real-time three.js pathtracer — a trained neural denoiser v0.4 plus multi samples-per-pixel frame support, moving from meshlets back to regular meshes, with moving lights, glass, caustics, materials, skinned meshes, morph targets, and global illumination presented as working in normal scenes. The denoiser is described as a small neural net with edge-aware filtering and per-pixel history at about 1.2 ms per frame, against OIDN post-processing near a second. No repository is public; the author has said he will decide on opening up after live demos.
Meaning: the budget tells the story. A denoiser that costs 1.2 ms inside the frame is part of the renderer; one that costs a second is part of the export pipeline. If the number holds, it is what makes the rest of the milestone possible. Evidence is currently one author’s posts plus demo video, with thin engagement and no code or independent measurement, so this sits at demonstrated-by-author rather than replicated.