From da19fce548c6b00d2d546783181946adb4458c55 Mon Sep 17 00:00:00 2001 From: Adam Ferguson Date: Sun, 6 Sep 2026 21:40:25 -0400 Subject: [PATCH 1/2] blog: painting in the Spark's spare memory (image-gen on the lab) --- src/content/blog/spark-image-generation.md | 87 ++++++++++++++++++++++ 1 file changed, 87 insertions(+) create mode 100644 src/content/blog/spark-image-generation.md diff --git a/src/content/blog/spark-image-generation.md b/src/content/blog/spark-image-generation.md new file mode 100644 index 0000000..291b8c9 --- /dev/null +++ b/src/content/blog/spark-image-generation.md @@ -0,0 +1,87 @@ +--- +title: "Painting in the Spark's spare memory" +description: "Adding local image generation to the LLM lab — why ComfyUI lives outside the converge pass, why SDXL and not Flux on 121 GiB of unified memory, and what I learned about Grace-Blackwell's quirks along the way." +date: 2026-09-06 +tags: + - llm + - infrastructure + - self-hosting + - gpu + - image-generation +category: Infrastructure +--- + +The LLM lab has been humming for a few weeks now — one config file, one command, +a model that answers when my tools call. But there's a second GPU workload that's +been bugging me: image generation. Not for anything grand. I wanted a new profile +picture, and I wanted the machine that already thinks to also try its hand at +drawing. + +Turns out that's mostly a memory-etiquette problem, not a software problem. + +## The co-tenancy problem + +The Spark is a Grace-Blackwell box: 121 GiB of *unified* memory. The CPU and the +GPU don't have separate pools — one allocation budget, both spenders. My LLM +runtime spends most of it: weights resident, KV cache filling what's left, and a +carefully-tuned reserve keeping the host from swapping. + +So ComfyUI isn't installing into this house; it's moving into a room I have to +vacate every time it wants to paint. That reframing drove every decision: + +- **The model choice is a budget choice.** Flux looks stunning on a Spark demo, + but at ~23 GB of weights plus a big text encoder, it only fits while nothing + else is loaded. SDXL base is ~7 GB — enough to generate a portrait in the + gaps between model work without evicting the LLM's KV pool. The fetch script + in the repo downloads exactly one file and says why in its header comment. +- **The service lives outside the convergence loop.** Everything else in the lab + is declarative: describe the state, `apply` makes it so. Image generation is + opt-in and explicit — `spark-lab comfyui up`, `down`, `status`, `logs` — + because `apply` re-converging a model host should never surprise a half-done + render, and a render I forgot about should never OOM a benchmark. The + scheduler and the diffusion model are adults; they get their own command. + +## Grace-Blackwell bites + +Two platform quirks worth writing down before I forget them: + +**Dynamic VRAM misreads unified memory.** ComfyUI's dynamic model-loading mode +asks CUDA how much VRAM is free and behaves like a normal discrete GPU box. On +unified memory, that answer reflects what CUDA thinks the process may use, not +what the machine has — and the resident LLM looks like "no room," so ComfyUI +offloads aggressively and doubles its own generation time for nothing. The +community fix is `--disable-dynamic-vram` plus a small `--reserve-vram`: tell +it the models stay resident and stop second-guessing. The default render ships +those flags. + +**The installer is a trust decision, not a convenience one.** Several Spark +ComfyUI images exist. I picked the one maintained by a ComfyUI org maintainer +(CUDA 13.1, sm_121, SageAttention, non-root), because the interesting Spark +setups are build-your-own — which is fine for a hobby afternoon and less fine +as infrastructure I want `status` to tell the truth about. It's overridable +through the same `images:` map everything else in the lab uses, so if I change +my mind it's a config edit. + +## The portrait, or: prompts are palettes + +The workflow files committed with this feature generate stylized portraits from +a source photo via img2img — `VAEEncode` the face in, denoise around 0.6, and +structure survives while style floods in. No custom nodes, no face-ID adapters +that break on the next ComfyUI release: core nodes only, so the graph keeps +loading in two years. + +Two styles, both derived from this site's own palette, which I extracted from +computed styles rather than eyeballing: espresso ink `#15140F`-ish, cream +`#ECE6DC`, taupe `#A39A8C`, and one ember accent `#F0824A`. First is a two-tone +woodcut relief — carved strokes, flat planes, one burnt-orange ball as the nod +to the header. Second is a flat-vector headshot in four colors, the kind of +thing that survives being cropped to a 40-pixel circle next to a reply. + +The woodcut wins, if it behaves; we'll see what the actual renders look like +before this becomes a profile picture. The nice thing is the whole loop — +photo in, prompt tuned to a palette I already trust, output on disk in +`basedir/output/` — now runs on the same box that writes these posts' code. + +All of it is in the [spark-lab repo](https://github.com/AdamFerguson/spark-lab), +MIT as ever: the comfyui command, the compose template, the model-fetch script, +and both workflow JSONs under `workflows/`. From 49d2d89ee3da0e907be9fd155a0b7c9ef4d81ed5 Mon Sep 17 00:00:00 2001 From: Adam Ferguson Date: Sun, 6 Sep 2026 22:18:05 -0400 Subject: [PATCH 2/2] blog: rewrite - cut the padding, add the collaboration angle --- src/content/blog/spark-image-generation.md | 160 +++++++++++---------- 1 file changed, 85 insertions(+), 75 deletions(-) diff --git a/src/content/blog/spark-image-generation.md b/src/content/blog/spark-image-generation.md index 291b8c9..ea8baf5 100644 --- a/src/content/blog/spark-image-generation.md +++ b/src/content/blog/spark-image-generation.md @@ -1,6 +1,6 @@ --- -title: "Painting in the Spark's spare memory" -description: "Adding local image generation to the LLM lab — why ComfyUI lives outside the converge pass, why SDXL and not Flux on 121 GiB of unified memory, and what I learned about Grace-Blackwell's quirks along the way." +title: "ComfyUI on the lab, mostly over chat" +description: "Adding local image generation to the LLM lab, and what it was like to steer three parallel work streams from a phone: the memory budget decisions, the Grace-Blackwell quirks, and the parts the agent got wrong." date: 2026-09-06 tags: - llm @@ -11,77 +11,87 @@ tags: category: Infrastructure --- -The LLM lab has been humming for a few weeks now — one config file, one command, -a model that answers when my tools call. But there's a second GPU workload that's -been bugging me: image generation. Not for anything grand. I wanted a new profile -picture, and I wanted the machine that already thinks to also try its hand at -drawing. - -Turns out that's mostly a memory-etiquette problem, not a software problem. - -## The co-tenancy problem - -The Spark is a Grace-Blackwell box: 121 GiB of *unified* memory. The CPU and the -GPU don't have separate pools — one allocation budget, both spenders. My LLM -runtime spends most of it: weights resident, KV cache filling what's left, and a -carefully-tuned reserve keeping the host from swapping. - -So ComfyUI isn't installing into this house; it's moving into a room I have to -vacate every time it wants to paint. That reframing drove every decision: - -- **The model choice is a budget choice.** Flux looks stunning on a Spark demo, - but at ~23 GB of weights plus a big text encoder, it only fits while nothing - else is loaded. SDXL base is ~7 GB — enough to generate a portrait in the - gaps between model work without evicting the LLM's KV pool. The fetch script - in the repo downloads exactly one file and says why in its header comment. -- **The service lives outside the convergence loop.** Everything else in the lab - is declarative: describe the state, `apply` makes it so. Image generation is - opt-in and explicit — `spark-lab comfyui up`, `down`, `status`, `logs` — - because `apply` re-converging a model host should never surprise a half-done - render, and a render I forgot about should never OOM a benchmark. The - scheduler and the diffusion model are adults; they get their own command. - -## Grace-Blackwell bites - -Two platform quirks worth writing down before I forget them: - -**Dynamic VRAM misreads unified memory.** ComfyUI's dynamic model-loading mode -asks CUDA how much VRAM is free and behaves like a normal discrete GPU box. On -unified memory, that answer reflects what CUDA thinks the process may use, not -what the machine has — and the resident LLM looks like "no room," so ComfyUI -offloads aggressively and doubles its own generation time for nothing. The -community fix is `--disable-dynamic-vram` plus a small `--reserve-vram`: tell -it the models stay resident and stop second-guessing. The default render ships -those flags. - -**The installer is a trust decision, not a convenience one.** Several Spark -ComfyUI images exist. I picked the one maintained by a ComfyUI org maintainer -(CUDA 13.1, sm_121, SageAttention, non-root), because the interesting Spark -setups are build-your-own — which is fine for a hobby afternoon and less fine -as infrastructure I want `status` to tell the truth about. It's overridable -through the same `images:` map everything else in the lab uses, so if I change -my mind it's a config edit. - -## The portrait, or: prompts are palettes - -The workflow files committed with this feature generate stylized portraits from -a source photo via img2img — `VAEEncode` the face in, denoise around 0.6, and -structure survives while style floods in. No custom nodes, no face-ID adapters -that break on the next ComfyUI release: core nodes only, so the graph keeps -loading in two years. - -Two styles, both derived from this site's own palette, which I extracted from -computed styles rather than eyeballing: espresso ink `#15140F`-ish, cream -`#ECE6DC`, taupe `#A39A8C`, and one ember accent `#F0824A`. First is a two-tone -woodcut relief — carved strokes, flat planes, one burnt-orange ball as the nod -to the header. Second is a flat-vector headshot in four colors, the kind of -thing that survives being cropped to a 40-pixel circle next to a reply. - -The woodcut wins, if it behaves; we'll see what the actual renders look like -before this becomes a profile picture. The nice thing is the whole loop — -photo in, prompt tuned to a palette I already trust, output on disk in -`basedir/output/` — now runs on the same box that writes these posts' code. - -All of it is in the [spark-lab repo](https://github.com/AdamFerguson/spark-lab), +The lab has run for a few weeks now: one config file, one command, a model that +answers when my tools call. The next workload to fit was image generation. Not +for anything grand. I wanted a new profile picture and I wanted the machine +that already thinks to try drawing one. + +Most of what I learned along the way was about memory, not software. + +## One pool, two spenders + +The Spark is a Grace-Blackwell box with 121 GiB of unified memory. CPU and GPU +share it. My LLM runtime spends most of it: weights resident, KV cache filling +what's left, a tuned reserve keeping the host off the swap. A diffusion model +moving in doesn't install, it moves into a room I have to vacate while it paints. + +Three consequences: + +The model choice is a budget choice. Flux demos beautifully on a Spark, but at +about 23 GB of weights plus its text encoder it only fits while nothing else is +loaded. SDXL base is about 7 GB, enough to generate beside a loaded LLM. The +fetch script in the repo downloads one file and its header says why. + +The service lives outside the convergence loop. Everything else in the lab is +declarative: describe the state, `apply` makes it so. Image generation didn't +want that. A gateway restart should never surprise a half-finished render, and a +render I forgot about should never OOM a benchmark. So `apply` never touches +ComfyUI. It gets its own command: `spark-lab comfyui up|down|status|logs`. + +The installer is a trust decision. Spark ComfyUI images exist at several levels +of homemade. I went with one maintained by a ComfyUI org maintainer (CUDA 13.1, +sm_121, SageAttention, non-root), overridable through the same `images:` map +everything else uses, so a change of heart is a config edit. + +## Two Grace-Blackwell bites + +ComfyUI's dynamic VRAM mode asks CUDA how much memory is free and behaves like +the box has a discrete GPU. On unified memory the answer reflects what CUDA +thinks the process may use, not what the machine has, so the resident LLM reads +as "no room" and ComfyUI offloads aggressively. Twice the generation time for +nothing. The community fix: `--disable-dynamic-vram --reserve-vram 4`. The +default render ships those flags. + +Worth naming the other thing I learned the hard way: agents will confidently +tell you wrong stuff. I asked my setup whether Matrix encryption was feasible +and it described a stack that wasn't what it runs. Right answer required reading +the actual adapter, which showed E2EE is supported today and only needs a +dependency that no longer compiles on modern macOS. When the agent says "X is +impossible," check what it actually read. + +## The collaboration part, which was the surprise + +This whole feature arrived over chat. Matrix, my phone, while I did other +things: I'd ask for something, the agent would work, I'd course-correct when a +stream went wrong. It wasn't one conversation. It was three streams at once: +staging a model recipe on the second machine, the ComfyUI integration as a +repo PR, and this post as another. Each had its own files and its own mistakes +to catch. + +The checkpoints that mattered were decisions, not keystrokes. Approving or +refusing a risky install step. Killing an experiment I hadn't consented to. +Telling the agent my schedule ("free the box tonight") and having it queue work +against that instead of fighting for memory now. Deciding that SDXL over Flux +was the right call for a machine that has a day job. Reviewing the PR text and +saying the blog draft reads too much like a blog written by a robot. + +What I didn't do: babysit. The work happened in parallel, persisted across +sessions, and showed up in reviewable units (a branch, a PR, a post). What the +agent got wrong it also, eventually, flagged itself: the workflows here ship +with core nodes only and a test plan that admits the first render hasn't +happened yet. I'd rather review honest scaffolding than confident fiction. + +## The portrait + +Two workflows committed with the feature, both derived from this site's palette +extracted from computed styles: espresso ink, cream, taupe, one ember accent +`#F0824A`. A two-tone woodcut relief with a burnt-orange ball, and a flat-vector +headshot in four colors that survives being cropped to a 40-pixel circle. The +woodcut is the plan. Identity carries through img2img at denoise 0.6, so no +custom nodes and nothing to rot on the next ComfyUI release. + +The woodcut wins if the renders behave. We'll see. + +Everything is in the [spark-lab repo](https://github.com/AdamFerguson/spark-lab), MIT as ever: the comfyui command, the compose template, the model-fetch script, -and both workflow JSONs under `workflows/`. +and both workflows under `workflows/`.