<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:base="https://benlive.tv"><title>Ben Live Blog</title><id>https://benlive.tv/blog/</id><link href="https://benlive.tv/blog/feed.xml" rel="self"/><link href="https://benlive.tv/blog/"/><author><name>Ben McNulty</name></author><updated>2026-09-26T12:00:00-04:00</updated><entry><title>Rebuilding my WebXR portfolio, one headset test at a time</title><id>https://benlive.tv/blog/rebuilding-my-webxr-portfolio/</id><link href="https://benlive.tv/blog/rebuilding-my-webxr-portfolio/"/><updated>2026-09-26T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Rebuilding my WebXR portfolio, one headset test at a time&lt;/h1&gt;
&lt;p&gt;My WebXR portfolio had become a time capsule. The client work was still worth showing. So were the experiments with games, broadcasts, 360-degree media, and spatial interfaces. But the room holding all of it had fallen behind the rest of my site. Some of the writing was out of date. The materials looked like placeholders. Even the little “About this site” sign was an image of text that had been edited to change the copyright years.&lt;/p&gt;
&lt;p&gt;I wanted to keep the room and the history in it. I also wanted someone considering me for a job to be able to walk through it and find a current, cared-for piece of software.&lt;/p&gt;
&lt;p&gt;I worked through the refresh with GPT-6 in Codex. The useful part of that collaboration was the iteration: inspect the existing implementation, make a bounded change, test it, then put on the headset and see what the tests missed. My first Quest 3 pass looked much better and felt smooth. It also gave me a very specific list of things to fix.&lt;/p&gt;
&lt;h2&gt;Keeping the gallery, replacing the machinery&lt;/h2&gt;
&lt;p&gt;The original scene used A-Frame. That was a useful way to build the first version: describe the room in HTML, attach components, and get into a headset quickly. Over time, the page accumulated dependencies and behaviors that were harder to inspect together.&lt;/p&gt;
&lt;p&gt;The refreshed gallery uses Three.js directly, with WebXR for the headset session and cannon-es for device physics. The exhibits, positions, text, and material choices live in a scene manifest. Loading, rendering, navigation, grabbing, and session management have separate modules.&lt;/p&gt;
&lt;p&gt;That separation matters when something goes wrong. If a tablet moves strangely, I can inspect its body and the hand holding it. If a sign is hard to read, I can change its lettering without rewriting the movement code. The project images and arrangement still tell the same story.&lt;/p&gt;
&lt;p&gt;The comparisons in this post use the same camera position, direction, field of view, and image dimensions. I captured the old version from the live site before replacing it. The refreshed views come from the local build. These are browser captures, not headset recordings; navigation overlays are hidden in both versions so the room is easier to compare. Tap an image to inspect it at full size.&lt;/p&gt;
&lt;div class=&quot;table-scroll-wrapper&quot;&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Original live gallery&lt;/th&gt;&lt;th&gt;Refreshed gallery&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/entrance-before.webp&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/entrance-before.webp&quot; alt=&quot;Original entrance with flat colors and the large portfolio sign.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/entrance-after.webp?v=1b9f0cfa458e&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/entrance-after.webp?v=1b9f0cfa458e&quot; alt=&quot;The same entrance with finished materials, framed signage and colored heading backlights.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;Lighting that helps me read&lt;/h2&gt;
&lt;p&gt;The first lighting pass made the room feel more coherent, but it also made parts of it drab. Brightening everything would have brought back the washed-out look I was trying to leave behind.&lt;/p&gt;
&lt;p&gt;The more useful change was to decide what should produce light and what should receive it. The description panels now read as light boxes. Photographs have illuminated frames. Large section headings keep a crisp white face with colored light inset behind the letters. The room has warmer rugs, richer teal timber, mineral walls, and a few restrained decorative details. At the entrance, the big WebXR script is now a rich violet against a softly graded teal enamel panel. The saturated lettering and softer backdrop make that first view feel more welcoming, while a dark edge keeps the script readable where it crosses the large portfolio letters.&lt;/p&gt;
&lt;p&gt;Looking at the comparison images also reminded me of something the original got right: those green-to-purple dividers. The first refresh had turned them a fairly lifeless gray. I brought the gradient back with a brighter green and a stronger violet. The color is part of the surface, so it still receives the room&apos;s lighting. A more realistic gallery should still feel like a joyful place to visit.&lt;/p&gt;
&lt;p&gt;One last headset pass caught flickering where the script letters overlap. The transparent padding around those letters was writing competing depths at the same surface. Keeping it out of the depth buffer, while still checking the lettering against the solid room, fixed that without adding another rendering pass.&lt;/p&gt;
&lt;p&gt;Most of that effect does not require another light in the scene. The panels use luminous materials, the local halos are small pieces of geometry, and the lettering gets its edge treatment in the text shader. Two actual feature lights give the room some direction. There is no real-time shadow pass for every frame on the wall.&lt;/p&gt;
&lt;p&gt;The surface detail is shared, too. Plaster, ash, woven fabric, terrazzo, powder-coated metal, and aluminium use small textures generated once when the gallery loads. Their scale is based on metres in the room, so stretching a wall does not turn a fine texture into an enormous pattern. Repeated static pieces are combined where that saves drawing work without making the whole building one indivisible object.&lt;/p&gt;
&lt;div class=&quot;table-scroll-wrapper&quot;&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Original live gallery&lt;/th&gt;&lt;th&gt;Refreshed gallery&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/gallery-before.webp&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/gallery-before.webp&quot; alt=&quot;Original view of the right-hand section dividers.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/gallery-after.webp?v=ab7d1158fed4&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/gallery-after.webp?v=ab7d1158fed4&quot; alt=&quot;Matching divider view with colored heading backlights, framed exhibits and richer materials.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;A hand should behave like a hand&lt;/h2&gt;
&lt;p&gt;The headset exposed a problem that looked fine with controllers. The hand models lined up when I held the controllers, then twisted at the wrist when I put them down. The controller&apos;s pointing direction was being used for a job that needed the hand&apos;s actual anatomy.&lt;/p&gt;
&lt;p&gt;The new gloves follow the wrist and finger joints supplied by WebXR. A controller ray still does the pointing; it no longer decides how an articulated wrist should be rotated. Both input modes now wear the same teal fabric glove, with softly padded palms, continuous fingers, and a small pale trim detail. The controller version uses the existing button-driven animation skeleton to pose that glove. Switching input modes no longer means switching to a completely different-looking hand.&lt;/p&gt;
&lt;p&gt;My next observation was about grabbing. A precise thumb-and-forefinger pinch worked, but that is not how I naturally pick up a phone. I wanted to curl my fingers around it or support it on an upward-facing palm.&lt;/p&gt;
&lt;p&gt;The interaction now checks the palm and fingers against the device&apos;s oriented collider. A curled grip can acquire it without a pinch. Opening that grip releases it. An open palm can support it, and tilting the palm lets it go. Once a grip is established, the device keeps its position relative to the hand instead of snapping to an arbitrary pose. Release returns it to physics with the recent motion of the hand.&lt;/p&gt;
&lt;p&gt;Passing a device between hands exposed another small but uncomfortable problem. Both hands cannot own the same phone. A successful second grab now transfers it to that hand, keeping its position and rotation at the moment of transfer. The first hand can open or move away without dropping it, and fingers that are still curled do not immediately grab it back.&lt;/p&gt;
&lt;p&gt;This is still an interaction model, not a simulation of skin and pressure. It needs to tolerate tracking noise and small occlusions without making a tablet jump. The important distinction is that the gesture follows the thing I am trying to do.&lt;/p&gt;
&lt;p&gt;Pinching away from an object still activates the teleport arc. While holding that pinch, twisting the wrist changes the arrow at the destination, giving hands a counterpart to the controller thumbstick&apos;s landing-direction control. A small dead zone and smoothing keep tiny wrist movements from making the arrow twitch.&lt;/p&gt;
&lt;p&gt;A later headset pass found a more disorienting issue: a brief interruption in pinch recognition could send me to the destination before I meant to move. The release now needs a short, sustained opening of the thumb and finger. An outer ring fills during that confirmation, with the destination and facing held still. Pinching again resumes the aim. Losing tracking cancels it. WebXR also distinguishes a completed selection from a cancelled one; treating every end event as permission to travel was too trusting.&lt;/p&gt;
&lt;p&gt;When I stopped to read a tablet or inspect the AR model, tiny movements in my hands made the held content jitter. I added a short settling period: once the hold is quiet, small fluctuations are softened; deliberate movement brings it straight back to following my hand. Position, rotation and scale have bounded tolerances, and the smoothing uses elapsed time so it behaves consistently at different refresh rates. The hands and their interaction checks still use the original tracking data.&lt;/p&gt;
&lt;h2&gt;Current content, without erasing the old projects&lt;/h2&gt;
&lt;p&gt;A portfolio can be technically current and still give someone the wrong impression if its content is stale.&lt;/p&gt;
&lt;p&gt;The résumé posters and pickup tablets now use captures of my actual &lt;a href=&quot;/about/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;About page&lt;/a&gt;. A small capture script arranges complete sections onto two sheets, checks that they fit, and records what was captured. The same images are used on the wall and on the tablets, with matching proportions. They also contain fewer pixels than the old square textures, despite presenting more current information.&lt;/p&gt;
&lt;p&gt;The tablet backs say “Ben McNulty.” The contact panel points to LinkedIn for professional conversations and &lt;code&gt;@benlivenow&lt;/code&gt; on Instagram and Threads for personal updates. The personal panel says my partner and I have been together since 2012, which will remain true without an annual edit. Descriptions of completed projects stay in the past tense.&lt;/p&gt;
&lt;div class=&quot;table-scroll-wrapper&quot;&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Original live gallery&lt;/th&gt;&lt;th&gt;Refreshed gallery&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/resume-before.webp&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/resume-before.webp&quot; alt=&quot;Old résumé posters on the gallery wall.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/resume-after.webp?v=1e04e7f8b08c&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/resume-after.webp?v=1e04e7f8b08c&quot; alt=&quot;New résumé posters captured from the current About page, with clear sightlines.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The social podium now carries extruded Instagram and Threads marks and a readable, luminous handle. Those shapes are built from the official brand vectors. They are geometry in the scene, with a little depth and a material finish, rather than another screenshot pretending to be a physical object.&lt;/p&gt;
&lt;div class=&quot;table-scroll-wrapper&quot;&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Original live gallery&lt;/th&gt;&lt;th&gt;Refreshed gallery&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/podium-before.webp&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/podium-before.webp&quot; alt=&quot;Original social podium with platform names written as text.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/podium-after.webp?v=e8474756b6fb&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/podium-after.webp?v=e8474756b6fb&quot; alt=&quot;The refreshed podium with extruded logos, illuminated handle and updated résumé tablets.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The credits sign is maintained as text now. It names the current technology, credits Claude, Codex, and Muse for the enhancement work, and keeps the “made with love” sentiment with a small red heart, rounded out from the wall with a flat back. I asked for another pass on that shape: a plump little sculpture felt more fitting than a beveled cutout. Updating a sentence no longer means editing an image of a sentence.&lt;/p&gt;
&lt;div class=&quot;table-scroll-wrapper&quot;&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Original live gallery&lt;/th&gt;&lt;th&gt;Refreshed gallery&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/credits-before.webp&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/credits-before.webp&quot; alt=&quot;The old credits image listing the original tools and copyright years.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/credits-after.webp?v=7c652b1cd979&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/credits-after.webp?v=7c652b1cd979&quot; alt=&quot;Editable credits with the current stack, GPT-6 attribution and a three-dimensional heart.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;Stepping inside a photograph&lt;/h2&gt;
&lt;p&gt;The 360-degree photos were already in the gallery, wrapped around little spheres. I wanted picking one up to mean something. They now sit in illuminated cradles, with smoother surfaces and a restrained crystal finish. Two small, light-faced signs explain the interaction before I reach the spheres, one on each path; the individual pedestal tops have their own colored inlays instead of repeating the same instructions. Holding one starts a short countdown. A small viewing platform rises under me and its rail extends before the gallery fades into the photograph.&lt;/p&gt;
&lt;p&gt;The important part is what stays still. Moving the ball in my hand does not move the photograph, turn my head, or carry the platform around. I can look over the rail and down into the image while keeping a familiar floor beneath me. Releasing the ball brings the gallery back with another fade, and the ball follows a soft curve home to its stand. Passing it to my other hand keeps the same visit going. A tracking interruption now freezes the held crystal instead of ending the visit. Returning requires a brief, continuous open-hand gesture, with a small progress ring confirming it. If tracking comes back after my hand has moved, the crystal resumes from its last position without jumping. Desktop visitors can click a crystal and click again, or press Escape, to return.&lt;/p&gt;
&lt;p&gt;The crystal keeps its small display texture, while the surrounding view loads a separate viewing version on demand. Most of my original files are only 1024 by 512 pixels. Wrapped Lanczos resampling and a light sharpening pass make the enlargement smoother without inventing detail. A two-image cache limits the extra texture memory; finding the original high-resolution files can come later. A small rim shader and 36 points provide the light inside the held ball; there is no second scene being rendered through it. These are panoramic photographs, so they surround the viewer without pretending to contain walkable three-dimensional geometry. The rail explains that limit, and the navigation controls pause for the visit.&lt;/p&gt;
&lt;div class=&quot;table-scroll-wrapper&quot;&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;On display in the gallery&lt;/th&gt;&lt;th&gt;Inside the photograph&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/crystal-display.webp?v=04a6577ed4b3&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/crystal-display.webp?v=04a6577ed4b3&quot; alt=&quot;Smooth photographic crystals on illuminated gallery stands.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;/img/blog/webxr-refresh/platform.webp?v=7e63bebe60a3&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/platform.webp?v=7e63bebe60a3&quot; alt=&quot;A small viewing platform and rail surrounded by the crystal&apos;s panoramic photograph.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;I also wanted the room to sound inhabited. The first soundscape was too static: long held tones felt more like a drone than a room I wanted to spend time in. It now moves through a little over four minutes of soft chord swells and sparse bell-like phrases, with changes in register, stereo placement and breathing space. Quiet answering notes and small changes in tone give the phrases more detail without adding more voices. The two entry corners now have ceramic and metal water sculptures, with water sounds positioned at the basins. I can hear the stream clearly up close, and it falls away as I walk farther into the room. Moving ripples and downward-flowing highlights give the water some life using the existing surfaces, with no extra lights or particle systems. Small interaction tones acknowledge a pickup, countdown, return, or teleport. Sound starts when I enter VR or AR. An embossed speaker button on the entrance wall or AR pedestal makes it easy to mute; M on a keyboard and B/Y on a controller still work. An explicit mute is remembered. The same musical bed continues in AR. There, the water sounds move to the room-scale sculpture accents and play more quietly, instead of staying at the old gallery coordinates.&lt;/p&gt;
&lt;h2&gt;AR needs its own way of looking at the room&lt;/h2&gt;
&lt;p&gt;Entering AR should give me a useful view of the gallery in my physical space. A life-size room dropped around me is not particularly helpful for that.&lt;/p&gt;
&lt;p&gt;The new AR mode presents the gallery as a miniature. The roof and overhead structure are removed from that view, so I can look down into it. Four mint-colored grip rails, mounted on the outside walls, control movement. One handle lets me lift and turn it in any direction; two handles let me change its size and facing. The grip points stay attached to my hands as I turn my wrists. Rotating the model around its own center looked plausible on a screen, but in the headset it pulled the handle away from my palm. The transform now follows the hand around the point where I grabbed it. With two hands, both grip points determine the position and size, while the wrists guide its tilt. The rails accept a palm-down grip as well as a pinch: the acquisition check includes my palm and curled knuckles, with a brief allowance for closing my hand just before reaching the rail. Touching the walls or reaching toward a tablet no longer grabs the entire room. It starts below eye level at a smaller, easier-to-survey scale, facing the entrance, using the headset’s current viewing pose. That makes it feel intentionally presented instead of left on the floor. The hanging feature-light fixtures are left out along with the roof.&lt;/p&gt;
&lt;p&gt;The miniature rests on a low pedestal that stays put when I lift the gallery away. It can float wherever I leave it. When the gallery overlaps the pedestal, or comes within its small magnetic margin, the rim changes color and asks me to release. The collision check follows the rotated gallery volume, so I don’t have to line up its center precisely. It then settles back into its original position, facing and size. Letting go far away leaves it alone, and losing hand tracking does not trigger a return.&lt;/p&gt;
&lt;p&gt;The larger welcome placard sits higher ahead of me, with a fountain on the floor at each side and a familiar rug underfoot. Supported surface detection checks for a stable floor and room boundaries. The sign and both fountains fit together as one arrangement in front of the detected walls; smaller rooms get a more compact version with all three features still present. The pedestal, rug and welcome arrangement share a placement plan. Entry waits for a stable headset pose, and the pedestal starts farther forward with room left for me to stand. Floor fitting can move it a short distance to find a supported footprint. Once I pick up the gallery, its pedestal stays put. Placement settles once, rather than chasing every tracking adjustment. Without that data, the local-floor reference provides a clearly labeled preview and a floating welcome arrangement; the controls include a rescan option. These are virtual furnishings, of course. The pedestal is not something to lean on.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;/img/blog/webxr-refresh/ar-model.webp?v=7bd55f5321c9&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/ar-model.webp?v=7bd55f5321c9&quot; alt=&quot;The AR gallery on a low docking pedestal, with mounted grip rails, a raised welcome placard and two larger fountains flanking it.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Desktop capture of the AR arrangement. In the headset, the real room shows through the surrounding space.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The miniature reuses the gallery&apos;s exhibit geometry and materials, with a compact beveled base beneath the room. The full-size physics world pauses while it is open. The miniature also gets its own batching: with the entire room in view, combining static labels, crystal shells, and repeated construction pieces saves considerably more drawing work than the full-size room’s smaller visibility groups. Returning from AR restores the regular gallery and my previous desktop position. That avoids scaling a physics simulation down, then hoping it will return to normal afterward.&lt;/p&gt;
&lt;p&gt;A tiny résumé is still a tiny résumé. Touching a tablet or crystal in the model brings out a separate copy at a comfortable viewing size, independent of how large I have made the gallery. Controllers can point and select them too, with a small aiming line showing the target. I can grab these copies, turn them over, pass them to the other hand, and leave them floating where I want to look at them. A separate Return button puts each one away. Up to four can stay open at once. While either hand holds a gallery rail, pokes and selections pause for both hands. A brief pause after letting go keeps a releasing finger from opening a tablet by accident. The actual tablet and its physics remain where they were.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;/img/blog/webxr-refresh/ar-exhibits.webp?v=2762259a794f&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;&lt;img src=&quot;/img/blog/webxr-refresh/ar-exhibits.webp?v=2762259a794f&quot; alt=&quot;A full-size résumé tablet and crystal float beside one another in front of the AR gallery, each with its own Return button.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Desktop capture of persistent AR exhibits. Release leaves each copy in place; its Return button puts it away.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Apple’s gaze-and-pinch input also needs its own handling. Its temporary pinch sources arrive through session events, and the hand’s grip pose is separate from the gaze ray. The implementation uses that grip pose to carry an object and, when a rotation is available, to steer the arrival direction. Source removal and cancellation never complete a stale teleport. Crystal visits remain in place through those interruptions; completed releases still need a short confirmation interval. That path has simulated-input coverage; a physical Vision Pro check is still ahead.&lt;/p&gt;
&lt;h2&gt;Smooth is something I have to check&lt;/h2&gt;
&lt;p&gt;The Quest 3 feedback on the earlier refreshed build was encouraging: the scene looked good and movement felt smooth. That was useful evidence, but it did not prove that every later change would be free.&lt;/p&gt;
&lt;p&gt;Angled edges were still sharper than I wanted. Enabling native antialiasing looked noticeably better when I tried it in the Quest 3, so that smoother view is the baseline now. The headset quality choices remain available, along with an off setting for comparison. I want to keep improving the appearance without giving up comfortable movement. A stack of full-screen effects is not worth it if the room stops feeling smooth.&lt;/p&gt;
&lt;p&gt;Desktop rendering has its own resolution budget. On the development Mac, the actual browser canvas kept a roughly 60 Hz cadence across five views with smoothing enabled, including a Retina-sized buffer capped at 4.2 million pixels. Those measurements include the running gallery; an offscreen rendering benchmark had made the desktop cost look much worse. Testing the path people actually use matters.&lt;/p&gt;
&lt;p&gt;The automated checks cover loading, text layout, texture use, collisions, controller behavior, hand gestures, teleporting, and AR entry and exit. Desktop GPU measurements help me spot expensive changes. The headset passes confirmed the improved locomotion and edge quality and helped refine the grabbing. Automated checks now cover passing a device between hands, including the first hand letting go afterward. I then tried the hand-to-hand transfer in the Quest 3, where it worked smoothly. The miniature AR interactions have automated coverage, with broader device testing still ahead. I tried the crystal visits and AR inspection in the Quest 3, then used that feedback to refine accidental grabs, tracking interruptions and the presentation around the miniature. Those latest refinements still need their own headset check. A desktop timing result or an emulated headset cannot settle how every interaction will feel on every device.&lt;/p&gt;
&lt;p&gt;That is also why I wanted the gallery back in the main paths through the site. It gives visitors a working project to inspect alongside the résumé and project descriptions. The &lt;a href=&quot;/hire-me/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Hire Me page&lt;/a&gt; now gives it a direct entrance, alongside the other ways to learn about my experience.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;/webxr/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Explore the gallery&lt;/a&gt; on a regular screen, or open the same page in a compatible headset. The résumé and conventional portfolio remain available if a walk through a 3D room is not how you want to browse.&lt;/p&gt;
&lt;h3&gt;Technical references&lt;/h3&gt;
&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;https://threejs.org/docs/pages/WebXRManager.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Three.js WebXRManager&lt;/a&gt;: controller, grip and hand spaces; headset rendering controls.&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://www.w3.org/TR/webxr-hand-input-1/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;WebXR Hand Input&lt;/a&gt;: articulated joint poses and anatomical axes.&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://www.w3.org/TR/webxr/#input-actions&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;WebXR input actions&lt;/a&gt;: successful selections and cancelled actions have different event sequences.&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://developers.meta.com/horizon/design/hands/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Meta hand interaction design&lt;/a&gt;: comfort, feedback, and accommodating the limits of tracking.&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://webkit.org/blog/15162/introducing-natural-input-for-webxr-in-apple-vision-pro/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;WebKit natural input for Apple Vision Pro&lt;/a&gt;: temporary pinch sources, session events, gaze rays, and grip poses.&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;https://www.meta.com/brand/resources/instagram/instagram-brand/&quot; target=&quot;&lt;em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Instagram brand resources&lt;/a&gt; and &lt;a href=&quot;https://www.meta.com/brand/resources/threads/&quot; target=&quot;&lt;/em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Threads brand resources&lt;/a&gt;: source geometry for the podium marks.&lt;/li&gt;&lt;/ul&gt;</content></entry><entry><title>From QA Automation to AI Quality Architecture</title><id>https://benlive.tv/blog/from-qa-automation-to-ai-quality-architecture/</id><link href="https://benlive.tv/blog/from-qa-automation-to-ai-quality-architecture/"/><updated>2026-08-28T12:00:00-04:00</updated><content type="html">&lt;h1&gt;From QA Automation to AI Quality Architecture&lt;/h1&gt;
&lt;p&gt;There&apos;s a pattern in AI engineering job postings that I don&apos;t think teams have noticed about themselves. They want people who can build with AI tools, and they&apos;re drowning in regressions because nobody on the team thinks like a tester.&lt;/p&gt;
&lt;p&gt;I didn&apos;t plan to end up here. I spent years breaking software and building the infrastructure that breaks it automatically, and then the tools changed underneath me. It turned out the instincts transferred almost perfectly.&lt;/p&gt;
&lt;h2&gt;What the tester&apos;s mindset actually is&lt;/h2&gt;
&lt;p&gt;It isn&apos;t pessimism. It&apos;s a habit of asking, for every claim of success, &quot;how would I know if that were false?&quot;&lt;/p&gt;
&lt;p&gt;An assistant says the tests pass. How would I know if the runner had swallowed a crash? A feature works in the demo. How would I know if it worked in the second browser? The rate limiter rejects the eleventh request. How would I know if it also rejected the first one from a different visitor sharing the same proxy?&lt;/p&gt;
&lt;p&gt;That question is the whole job, and it&apos;s the exact question AI-assisted development needs asked constantly, because assistants are confident, fast, and unbothered by their own mistakes.&lt;/p&gt;
&lt;h2&gt;Selenium to Playwright, and what carried over&lt;/h2&gt;
&lt;p&gt;My automation history runs through Selenium, and my day job today includes a migration from a VB.NET Selenium codebase to C# on an adapter pattern that supports both Selenium and Playwright drivers underneath the same test logic. The page object model, the separation of locators from assertions, the discipline of deterministic waits over sleeps, all of that predates AI and all of it still matters.&lt;/p&gt;
&lt;p&gt;On this site the suite is Playwright, fifty-seven spec files across Desktop Chrome, Desktop Firefox, and Mobile Safari. Layout consistency measured in pixels. Accessibility scans with axe-core. Contrast ratios computed in the browser. Chat behavior against a mocked API. Full-page visual baselines at several viewports. A canary of eighteen sentinel tests that runs first and decides whether the full batched run is worth the time.&lt;/p&gt;
&lt;h2&gt;The false green&lt;/h2&gt;
&lt;p&gt;The best example I have of the mindset paying off happened during a full review of this site this month.&lt;/p&gt;
&lt;p&gt;The batched test runner, which I&apos;d been trusting for months, ran each batch under a three-minute timeout and parsed pass and fail counts from the output. If a batch crashed or got killed by the timeout, there was no &quot;N failed&quot; line to parse, so the count was zero, and zero failures meant &quot;ok.&quot; The screenshot-heavy batch takes fifteen minutes. It had been killed at three minutes and reported green on every run since the runner was written.&lt;/p&gt;
&lt;p&gt;Nothing about that would have been caught by asking the assistant whether the tests passed. It was caught by asking how I&apos;d know if they hadn&apos;t. The fix was two lines: a non-zero exit code with no parsable summary is a failure. The follow-up was more interesting, because the batch that had never finished turned out to contain fifteen tests that couldn&apos;t pass on WebKit at 4K and a set of visual baselines that had drifted long ago.&lt;/p&gt;
&lt;p&gt;The same review found a contract suite that skipped itself whenever the emulator was slow to start, with a message that made it look intentional. Silent skips and false greens are the two failure modes I now look for first in any test infrastructure, mine included.&lt;/p&gt;
&lt;h2&gt;Test-first thinking against AI-generated regressions&lt;/h2&gt;
&lt;p&gt;The regressions assistants introduce have a shape. They&apos;re rarely dramatic. They&apos;re a locator that matches a second element after a nav change. A retry set copied from the server to the client that quietly retries permanent errors. A CSS comment with a stray &lt;code&gt;*/&lt;/code&gt; that swallows the next rule. Each one passes a glance and fails a test, which means the test has to exist before the glance.&lt;/p&gt;
&lt;p&gt;So the rule on this site is that every bug found by hand becomes an automated check. The suite grows in the shape of how things actually break. The assistant reads the suite as context and stops making the same mistake, which is the part that&apos;s new.&lt;/p&gt;
&lt;h2&gt;Page objects to prompt contracts&lt;/h2&gt;
&lt;p&gt;The transferable pattern I didn&apos;t expect: a page object and a prompt contract are the same abstraction. Both isolate a volatile surface behind a stable interface. A page object hides the DOM so the test logic survives a redesign. A prompt key hides the system prompt so the client survives a prompt rewrite, and a versioned contract with a strict response shape means the UI can assert on the output the same way a test asserts on a page.&lt;/p&gt;
&lt;p&gt;Once I saw that, the AI Lab apps got easier to design. Four Problem Solver stages, four contracts, four independently testable units.&lt;/p&gt;
&lt;h2&gt;Quality architecture, not quality assurance&lt;/h2&gt;
&lt;p&gt;The role I&apos;m actually doing now is designing systems so they stay reliable while assistants and people change them at speed. Layered tests that fail loudly. Runners that can&apos;t lie. Context files that keep the assistant inside the lines. Gates that apply equally to every contributor. That&apos;s architecture, and it&apos;s the part of the AI transition that the hiring conversation keeps underrating.&lt;/p&gt;
&lt;p&gt;If you spent a decade breaking things, you already have the hardest skill for this work. The rest is tooling.&lt;/p&gt;</content></entry><entry><title>Human in the Loop: Designing AI Products People Trust</title><id>https://benlive.tv/blog/human-in-the-loop-designing-ai-products-people-trust/</id><link href="https://benlive.tv/blog/human-in-the-loop-designing-ai-products-people-trust/"/><updated>2026-08-25T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Human in the Loop: Designing AI Products People Trust&lt;/h1&gt;
&lt;p&gt;The hardest design problem in an AI product isn&apos;t making the model smarter. It&apos;s making the person feel in control of it. I&apos;ve been circling that question since university, where my research was on user interface considerations for emerging input device technologies. The input device has changed. The question hasn&apos;t.&lt;/p&gt;
&lt;p&gt;&quot;Human in the Loop&quot; is the tagline on this site, so I should be able to say concretely what it means in the things I&apos;ve built. Here&apos;s what it means.&lt;/p&gt;
&lt;h2&gt;Every inference is explicit&lt;/h2&gt;
&lt;p&gt;In the AI Lab apps, nothing generates on its own. Problem Solver proposes candidate problems when you click, analyzes the one you chose when you click, proposes solutions when you click, and builds a plan when you click. Agent Agenda has four generation actions and every one is a button. Promptpad enhances a draft when you ask and shows you the diff before you keep it.&lt;/p&gt;
&lt;p&gt;That was a deliberate constraint, and it has costs. The apps are slower to use than they would be with speculative generation. But there is never a moment where output appears that you didn&apos;t ask for, and that turns out to matter more than speed. Agency is the thing people notice first when it&apos;s gone.&lt;/p&gt;
&lt;h2&gt;Approval is a checkpoint, not a formality&lt;/h2&gt;
&lt;p&gt;Sequential gating creates natural checkpoints, and the design makes them real. In Problem Solver you can go back to an earlier validated step without spending another inference. In Promptpad the committed revision is distinct from the active one, so you can explore and still return to the version you trusted. The person&apos;s judgment is stored state, not a modal you click through.&lt;/p&gt;
&lt;p&gt;The same pattern runs my development workflow. Claude Code writes a plan, I approve the plan, then it executes. &lt;code&gt;/diagnose&lt;/code&gt; produces hypotheses and stops before a fix. A checkpoint that requires a decision is a feature in a tool and a feature in a process.&lt;/p&gt;
&lt;h2&gt;Honesty about what happened&lt;/h2&gt;
&lt;p&gt;When the chat&apos;s Balanced model is unavailable and the fallback chain answers with a different one, the response says which model answered. The status badge and the reply label can disagree, and that&apos;s correct. I&apos;d rather a visitor see the truth than a tidy interface.&lt;/p&gt;
&lt;p&gt;The same goes for reasoning. The Thoughtful tier uses a model that produces reasoning traces, and the server preserves them for continuity. None of the three chats display them, because the prompts explicitly prohibit exposing structured reasoning to the visitor. Showing a model&apos;s scratchpad as if it were an explanation would be theater, and theater erodes trust faster than silence.&lt;/p&gt;
&lt;h2&gt;Consent that&apos;s real&lt;/h2&gt;
&lt;p&gt;Analytics on this site runs through a consent layer that&apos;s checked before anything fires. The chat panels on the Context First and prehog presentations route message text through a completely separate path from analytics, and a dedicated privacy test suite asserts that a typed message never appears in any captured event. Session replay, where it runs, masks every input and the entire chat dialog, and it&apos;s scoped to one page to answer one question.&lt;/p&gt;
&lt;p&gt;That consent architecture started on one presentation and became a small shared library used across the site, because building it properly once made it obvious what everywhere else was missing.&lt;/p&gt;
&lt;h2&gt;The &quot;would I use this&quot; test&lt;/h2&gt;
&lt;p&gt;Every AI feature I ship has to pass one question: would I use this myself, on a bad day, when I&apos;m skeptical? Auto-generated output that I&apos;d have to undo fails. Hidden model swaps fail. A &quot;thinking&quot; panel that&apos;s really a transcript dump fails. Analytics I can&apos;t turn off fail.&lt;/p&gt;
&lt;p&gt;What passes is boring in the best way. One click, one result, a clear label, a way back.&lt;/p&gt;
&lt;h2&gt;What enterprise buyers mean by responsible AI&lt;/h2&gt;
&lt;p&gt;At work, on a government software contract, &quot;responsible AI&quot; isn&apos;t a slogan. It&apos;s whether a reviewer can see what the assistant changed, whether the change ran through the same gates as any other change, and whether the person who approved it could explain why. Approval checkpoints, honest labels, scope control. It&apos;s the same list.&lt;/p&gt;
&lt;p&gt;I don&apos;t think this is a temporary phase before the models get good enough to be trusted unsupervised. I think it&apos;s the design. The loop is the product.&lt;/p&gt;</content></entry><entry><title>Context Engineering: The Skill That Makes AI Assistants Useful</title><id>https://benlive.tv/blog/context-engineering-the-skill-that-makes-ai-assistants-useful/</id><link href="https://benlive.tv/blog/context-engineering-the-skill-that-makes-ai-assistants-useful/"/><updated>2026-08-23T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Context Engineering: The Skill That Makes AI Assistants Useful&lt;/h1&gt;
&lt;p&gt;Most engineers who decide AI assistants &quot;aren&apos;t ready&quot; got there the same way. They opened a chat, typed a vague request into an empty context, got a mediocre answer, and drew a conclusion about the model. The model wasn&apos;t the variable. The context was.&lt;/p&gt;
&lt;p&gt;Context engineering is the discipline of shaping what an assistant knows before it generates anything. It&apos;s the difference between a tool that produces mergeable changes and an expensive autocomplete. I&apos;ve been doing it long enough now, across my own repos and a team codebase at work, to have opinions about what actually moves the needle.&lt;/p&gt;
&lt;h2&gt;Rule files are contracts, not notes&lt;/h2&gt;
&lt;p&gt;On this site, &lt;code&gt;CLAUDE.md&lt;/code&gt; is a contract. It says which runtime to use, how the two git repositories relate, that Playwright runs in batches and never as one bulk process, that the prismatic gradient lives in exactly one CSS file, that &lt;code&gt;transition: all&lt;/code&gt; is banned, and that a bug gets a written diagnosis before it gets a fix. When a new session starts, the assistant already knows all of that. I never paste it.&lt;/p&gt;
&lt;p&gt;The important part is that the file changes when the code changes. When I moved the model configuration into its own module, the rule file was updated in the same commit. When I found that the batch runner was hiding failures, the rule about running tests got more specific. A rule file that lags the code is worse than no rule file, because the assistant trusts it.&lt;/p&gt;
&lt;p&gt;At work the same idea lives as version-controlled Rules, Profiles, and Prompts that the whole team reviews. Different format, same contract.&lt;/p&gt;
&lt;h2&gt;Skills encode the steps people skip&lt;/h2&gt;
&lt;p&gt;Some parts of a good workflow are easy to state and hard to keep doing at the end of a long session. Those became Claude Code skills. &lt;code&gt;/diagnose&lt;/code&gt; refuses to fix anything until there&apos;s a ranked hypothesis list. &lt;code&gt;/implement-plan&lt;/code&gt; finds the existing plan and executes it instead of writing a fresh one. &lt;code&gt;/audit&lt;/code&gt; runs a checklist over the diff before a commit. A skill is context that arrives at the right moment instead of sitting in a file hoping to be remembered.&lt;/p&gt;
&lt;h2&gt;Server-side context control&lt;/h2&gt;
&lt;p&gt;Context engineering isn&apos;t only about what the assistant sees during development. It&apos;s also about what the model sees at runtime in a shipped feature.&lt;/p&gt;
&lt;p&gt;Every AI Lab app sends a &lt;code&gt;promptKey&lt;/code&gt;, never a prompt. The server resolves the key to a system prompt the visitor can&apos;t read or override. Each stage of Problem Solver has its own key with its own contract and a version suffix. That structure is what lets me change a prompt without changing the behavior of sessions created under the old one, and it&apos;s the reason a public chat can run without anyone being able to rewrite its instructions.&lt;/p&gt;
&lt;h2&gt;Memory that persists across sessions&lt;/h2&gt;
&lt;p&gt;The newest layer is a memory directory: one file per fact, with a short index loaded into every session. Not code structure, which the repo already records, and not past fixes, which git records. The things that aren&apos;t derivable: decisions I&apos;ve made about how work should be done, constraints that came from a conversation, pointers to resources outside the repo. It&apos;s the same principle as the rule file, scoped to what only a person would know.&lt;/p&gt;
&lt;h2&gt;Patterns that pay off immediately&lt;/h2&gt;
&lt;p&gt;Three habits account for most of the improvement I&apos;ve seen.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope with non-goals.&lt;/strong&gt; &quot;Change the rate limiter&apos;s memory behavior. Do not touch the Firestore path.&quot; The second sentence prevents more damage than the first one causes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exit criteria before starting.&lt;/strong&gt; What test passes when this is done? If I can&apos;t answer that, the task isn&apos;t specified yet.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Read the minimum.&lt;/strong&gt; The assistant reads the function and its callers, not the directory. Broad reads become broad edits.&lt;/p&gt;
&lt;h2&gt;Why this beats model selection&lt;/h2&gt;
&lt;p&gt;I&apos;ve watched a mid-tier model with a precise rule file and a tight scope out-deliver a top-tier model with neither. Model quality sets a ceiling. Context decides how close you get to it. And context is the part you control, on every task, for free.&lt;/p&gt;
&lt;h2&gt;Where it goes from here&lt;/h2&gt;
&lt;p&gt;I built &lt;a href=&quot;/context-first/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Context First&lt;/a&gt; as a presentation on exactly this idea, first as a job application and then generalized: context over control, decisions written down, small teams trusted with the final call. It&apos;s the way I want to work with people, and it turns out to be the way I have to work with assistants. If a company has written down how it operates somewhere I can check, I can be effective there on day one. The same is true of a codebase.&lt;/p&gt;</content></entry><entry><title>What a PostHog Application Taught Me</title><id>https://benlive.tv/blog/what-a-posthog-application-taught-me/</id><link href="https://benlive.tv/blog/what-a-posthog-application-taught-me/"/><updated>2026-08-20T12:00:00-04:00</updated><content type="html">&lt;h1&gt;What a PostHog Application Taught Me&lt;/h1&gt;
&lt;p&gt;I applied for the Context Engineer role on PostHog&apos;s Wizard &amp;amp; Docs team. I built the whole application as a working presentation instead of a cover letter: a real PostHog implementation, Product Analytics, masked Session Replay, a custom survey, a feature flag, an AI chat panel a reviewer could actually talk to. I didn&apos;t get the role.&lt;/p&gt;
&lt;p&gt;This is the honest version of what happened, what I&apos;d do differently, and what I&apos;m keeping from the process anyway.&lt;/p&gt;
&lt;h2&gt;The application&lt;/h2&gt;
&lt;p&gt;The idea was simple: instead of describing my skills, show them, on the exact product I was applying to work on. I built a nine-slide presentation at what&apos;s now &lt;code&gt;/prehog&lt;/code&gt;, instrumented it with PostHog&apos;s own SDK, and wrote up every decision, what I implemented and why, what I deliberately left out, in a public repository a technical reviewer could actually read.&lt;/p&gt;
&lt;p&gt;It worked as a demo. The mechanics held up: keyboard navigation, a no-JavaScript fallback, a reference mode for browsing instead of paging through slides, a live event log showing exactly what PostHog was and wasn&apos;t collecting in real time. I still think that part was the right call. If you&apos;re applying to build analytics tooling, shipping working analytics tooling says more than describing your approach to it.&lt;/p&gt;
&lt;h2&gt;What I got wrong&lt;/h2&gt;
&lt;p&gt;The writing is where I&apos;d change course. I leaned on AI assistance to draft and refine the narrative slides, the &quot;why PostHog&quot; case, the culture-fit argument, more than I should have. The result read as polished in a way that worked against me: competent, thorough, and a little too smooth. A more focused piece, written in a plainer voice with fewer claims doing more specific work, would have landed better.&lt;/p&gt;
&lt;p&gt;That&apos;s a strange thing to admit while applying to companies that build and rely on AI tooling. I don&apos;t think the answer is to avoid AI assistance. I think the answer is to use it for what it&apos;s actually good at, structure, research, catching gaps, and stay the one making the calls about what the piece actually says and how much of it needs saying. Somewhere in the drafting process I let the tool do too much of the thinking instead of the typing. The fix isn&apos;t less AI. It&apos;s a clearer line between assistance and authorship, and fewer, sharper examples of real work over more sentences describing it.&lt;/p&gt;
&lt;h2&gt;What stuck with me&lt;/h2&gt;
&lt;p&gt;Reading PostHog&apos;s own public handbook while I built the case for why I wanted the role turned into the more useful part of the process. Somewhere in there I picked up &lt;em&gt;No Rules Rules&lt;/em&gt;, Reed Hastings and Erin Meyer&apos;s account of Netflix&apos;s culture, because PostHog&apos;s own spending-money handbook cites it directly as the influence behind their context-based expense policy. It&apos;s also listed on PostHog&apos;s company story page under &quot;Things that influenced us.&quot; That&apos;s a company-level citation, well beyond the one HR policy.&lt;/p&gt;
&lt;p&gt;The book&apos;s actual argument, give people context instead of control, and trust them to act on it, showed up independently across PostHog&apos;s culture, communication, and small-teams handbooks too. Fully remote across 20-plus countries. Two weekdays kept free of internal meetings by default. A stated bias toward &quot;not asking for permission&quot; if you&apos;re acting in the company&apos;s interest. Small teams with the final call on what ships and no external QA gate second-guessing them. None of that reads as marketing copy. It reads as a company that wrote down how it actually wants to operate and then built the systems to make that true.&lt;/p&gt;
&lt;p&gt;I didn&apos;t get the role, but that pattern is worth more to me than one job outcome. It&apos;s the standard I&apos;m holding the rest of my search to now: not &quot;does this company say the right things,&quot; but &quot;has this company written its own operating principles down somewhere I can check, and do the systems actually match.&quot;&lt;/p&gt;
&lt;h2&gt;The technical part: building a real PostHog implementation&lt;/h2&gt;
&lt;p&gt;The application itself is still worth walking through, because the analytics work turned out to matter beyond the one role it was built for.&lt;/p&gt;
&lt;p&gt;I started with a short list of real questions instead of turning on blanket tracking: does anyone finish the presentation? Is the navigation discoverable on mobile? Do reviewers actually check the source code? Then I added exactly the events needed to answer each one, and nothing else. That question-first discipline is the part I&apos;d defend most as good engineering in its own right.&lt;/p&gt;
&lt;p&gt;A few pieces worth calling out specifically:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consent-gated Session Replay, masked by default.&lt;/strong&gt; Replay is scoped to one page, answers one question (is navigation discoverable on a phone), and masks every input plus the entire chat dialog as rendered text, extending coverage to the message bubbles the chat renders afterward, beyond the input field alone. PostHog&apos;s &lt;code&gt;maskAllInputs&lt;/code&gt; only covers that second case on its own.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A real Survey, not PostHog&apos;s default popover.&lt;/strong&gt; One rating question and one optional open-text field, custom-rendered in the page&apos;s own styling and shown once, after a visitor reaches the end. Restraint here felt more honest than a longer survey would have: one well-placed exchange beats a form nobody fills out.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A feature flag used for an actual rollout control&lt;/strong&gt;, the genuine kind. I&apos;d declined feature flags earlier in the build, since a flag with no real branch to test is just a fake experiment for résumé optics, then reversed the call once there was a genuine use for one: a self-referential live event-log panel worth being able to turn on for specific visitors without a redeploy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An AI chat panel with a hard privacy line.&lt;/strong&gt; Message text and model responses never reach PostHog at all; they route through a completely separate path (Firebase Functions to OpenRouter) with its own disclosure. The system prompt that drives it distinguishes explicitly between what&apos;s stated on the deck or independently verified and what&apos;s my own synthesis, and it&apos;s built to say plainly when it doesn&apos;t know something rather than guess.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A shared consent architecture that outlived the one page it started on.&lt;/strong&gt; Building this properly for one presentation exposed what a real analytics setup actually needs: a consent layer applied consistently everywhere, checked before anything fires. That generalized into a small library now used across this site&apos;s main public pages, carrying the same discipline everywhere: specific, meaningful events, with visitor control designed in from the start.&lt;/p&gt;
&lt;p&gt;None of that was wasted effort. It&apos;s still running, still instrumented, and still the clearest demonstration I have of how I actually build with PostHog&apos;s product rather than just talk about it.&lt;/p&gt;
&lt;h2&gt;Where it lives now&lt;/h2&gt;
&lt;p&gt;The original application stays exactly as it was, at &lt;a href=&quot;https://benlive.tv/prehog&quot; target=&quot;&lt;em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;benlive.tv/prehog&lt;/a&gt;, source at &lt;a href=&quot;https://github.com/benmcnulty/prehog&quot; target=&quot;&lt;/em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;github.com/benmcnulty/prehog&lt;/a&gt;, in case a future PostHog-specific opportunity ever makes it relevant again. I&apos;m not editing history to make it look like something it wasn&apos;t.&lt;/p&gt;
&lt;p&gt;What&apos;s live going forward is &lt;a href=&quot;https://benlive.tv/context-first&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;benlive.tv/context-first&lt;/a&gt;: the same working implementation, generalized. The narrative shifted from &quot;why hire me for this role&quot; to why this way of working, context over control, decisions written down, small teams trusted with the final call, is what I&apos;m actually looking for next, wherever that ends up.&lt;/p&gt;</content></entry><entry><title>Building Local Crew: A Local-First Inference Orchestration Framework</title><id>https://benlive.tv/blog/building-local-crew-port/</id><link href="https://benlive.tv/blog/building-local-crew-port/"/><updated>2026-07-13T12:00:00-04:00</updated><content type="html">&lt;p&gt;I built Local Crew because I wanted a system that could treat the machines already in my house like one coordinated AI environment instead of a pile of disconnected terminals.&lt;/p&gt;
&lt;p&gt;That&apos;s the core idea: take the laptop, desktop, mini PC, and whatever else you already have on your network, register them as resources, and let one orchestrator handle routing, delegation, task flow, memory, and quality checks across the whole group. No cloud account required. No API keys unless you want to add them. Your hardware, your models, your data.&lt;/p&gt;
&lt;p&gt;The source is on GitHub: &lt;strong&gt;&lt;a href=&quot;https://github.com/benmcnulty/localcrew&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;github.com/benmcnulty/localcrew&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Port matters here, and I&apos;ll cover it, but it isn&apos;t the main event. The main event is the local-first orchestration framework itself: a practical system for running real multi-device AI workflows on infrastructure you control.&lt;/p&gt;
&lt;h2&gt;What Local Crew Actually Is&lt;/h2&gt;
&lt;p&gt;Local Crew is a local-first orchestration harness for multi-device inference networks. The orchestrator sits in the middle, abstracts hardware behind named agent identities, and routes work to the best available resource based on capability, availability, memory headroom, and task complexity.&lt;/p&gt;
&lt;p&gt;In practice, that means one system can coordinate:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;a strong top-tier machine for long-context reasoning and delegation&lt;/li&gt;&lt;li&gt;mid-tier machines for the bulk of task execution&lt;/li&gt;&lt;li&gt;smaller devices for fast classification, routing, and utility work&lt;/li&gt;&lt;li&gt;optional OpenAI-compatible or Anthropic endpoints when you want them&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;The default path is still local inference. Ollama on your own network is the baseline. Everything else is additive.&lt;/p&gt;
&lt;p&gt;What makes the framework useful is that it doesn&apos;t stop at model access. It also gives you:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;autonomous queue filling and task execution through &lt;code&gt;/auto&lt;/code&gt;&lt;/li&gt;&lt;li&gt;persistent system memory and per-agent memory across sessions&lt;/li&gt;&lt;li&gt;resource discovery, topology management, and live capacity reporting&lt;/li&gt;&lt;li&gt;a browser dashboard at &lt;code&gt;/ui&lt;/code&gt; for direct operator control&lt;/li&gt;&lt;li&gt;a wallboard at &lt;code&gt;/display&lt;/code&gt; for always-on visibility&lt;/li&gt;&lt;li&gt;transactional audit and telemetry so the system stays legible as it grows&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;That&apos;s the difference between a model runner and an orchestration framework. Local Crew is designed to manage ongoing work, not just answer prompts.&lt;/p&gt;
&lt;h2&gt;The Operator Experience Matters&lt;/h2&gt;
&lt;p&gt;The CLI is still the center of gravity, because I want the full system to remain operable from the terminal. But I also wanted the state of the system to be visible without forcing everything through text.&lt;/p&gt;
&lt;p&gt;The local browser UI gives me a working control surface for resources, queue state, direct interactions, and configuration. It&apos;s the view I use when I want to inspect the system instead of talk to it.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/building-local-crew-port/localcrew-ui-dashboard.png&quot; alt=&quot;The Local Crew /ui dashboard showing resources, queue state, and cluster capacity.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;/ui&lt;/code&gt; dashboard is the practical operator surface for local control.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;/display&lt;/code&gt; wallboard solves a different problem. It&apos;s for the monitor across the room, the TV on the wall, or the always-on secondary display that keeps the whole system legible while work is running.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/building-local-crew-port/localcrew-display-wallboard.png&quot; alt=&quot;The Local Crew /display wallboard showing fleet status, current focus, metrics, and the live audit feed.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;/display&lt;/code&gt; wallboard turns the orchestrator into a visible live system instead of a hidden background process.&lt;/p&gt;
&lt;p&gt;That split is intentional. I want the framework to work as both a deep operator tool and a glanceable status surface. If I&apos;m going to trust a system with autonomous task execution, I need it to be inspectable.&lt;/p&gt;
&lt;h2&gt;Setup Is Intentionally Short&lt;/h2&gt;
&lt;p&gt;One of the main goals for Local Crew is to make the first useful run happen quickly.&lt;/p&gt;
&lt;p&gt;On the primary machine, the setup is straightforward:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;bash&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;git clone https://github.com/benmcnulty/localcrew.git
cd localcrew
npm install
npm run setup:crew&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That bootstrap flow discovers local models, writes the managed environment state, registers the primary resource, and starts the orchestrator in the same terminal.&lt;/p&gt;
&lt;p&gt;Adding another machine follows the same pattern:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;bash&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;npm run setup:agent&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The agent setup flow handles connection checks, device naming, model discovery, and sync back to the orchestrator. Once the network is online, the commands that matter are the ones you&apos;d expect:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;text&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;/status
/resource list
/topology
/auto&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That setup story is important to me because too many &quot;local AI&quot; projects fall apart between the demo and the second machine. I wanted this one to be useful as a real framework, not just interesting as a proof of concept.&lt;/p&gt;
&lt;h2&gt;How The Orchestration Loop Works&lt;/h2&gt;
&lt;p&gt;The orchestration model is where Local Crew stops being a thin wrapper around local models and starts becoming an actual system.&lt;/p&gt;
&lt;p&gt;Resources are organized into tiers. Top-tier devices can carry longer-context reasoning and even act as sub-orchestrators. Mid-tier devices handle the bulk of day-to-day work. Lower-tier devices can still be useful for fast lightweight tasks. The point isn&apos;t to pretend every machine is equal. The point is to use each one where it helps most.&lt;/p&gt;
&lt;p&gt;When Local Crew enters &lt;code&gt;/auto&lt;/code&gt;, it moves through a continuous loop:&lt;/p&gt;
&lt;ol&gt;&lt;li&gt;fill the queue through a draft-review-finalize consensus flow&lt;/li&gt;&lt;li&gt;route tasks to the most appropriate available resource&lt;/li&gt;&lt;li&gt;resolve tool use and file workflows as needed&lt;/li&gt;&lt;li&gt;verify completed work for substance and goal alignment&lt;/li&gt;&lt;li&gt;recover safely when something fails&lt;/li&gt;&lt;/ol&gt;
&lt;p&gt;That loop is what makes the framework feel different in practice. Instead of prompting one model over and over, you&apos;re operating a system that can generate work, distribute it, track it, and learn from the results over time.&lt;/p&gt;
&lt;p&gt;The agent layer matters here too. Participants have names, instructions, model bindings, and their own memory. That lets the orchestrator work with stable roles instead of raw endpoint strings, which makes the whole environment easier to reason about.&lt;/p&gt;
&lt;h2&gt;Port Is The Companion Layer, Not The Product&lt;/h2&gt;
&lt;p&gt;Port is the remote-facing layer that sits beside the local system, not above it.&lt;/p&gt;
&lt;p&gt;That&apos;s an important distinction. The local app stays fully functional on the LAN. &lt;code&gt;/ui&lt;/code&gt; and &lt;code&gt;/display&lt;/code&gt; are meant to be available locally out of the box. Remote authenticated access belongs in a separate surface, and that&apos;s where Port comes in.&lt;/p&gt;
&lt;p&gt;Port gives Local Crew a public-facing identity layer: sign-in, a Captain handle, pairing, and a remote dashboard that can eventually accept authenticated task submission from outside the local network.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/building-local-crew-port/port-captain-dashboard.png&quot; alt=&quot;The authenticated Port dashboard showing a connected Captain profile, live crew HUD, queue metrics, and recent tasks.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;p&gt;Port gives the orchestrator a remote identity and a public-facing dashboard without turning the local app into a cloud dependency.&lt;/p&gt;
&lt;p&gt;The auth stack is pragmatic. I used Firebase Auth with Google Sign-In because I wanted low-friction onboarding, sane session handling, and reuse of infrastructure I already trust in the rest of benlive.tv. After sign-in, the user claims a Captain handle and pairs a Local Crew instance through a short-lived device token entered into the CLI with &lt;code&gt;/login &amp;lt;token&amp;gt;&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;That pairing flow matters, but it only makes sense because the local framework already exists. Port is useful precisely because the real application lives on your own machines first.&lt;/p&gt;
&lt;h2&gt;Why I Think This Model Matters&lt;/h2&gt;
&lt;p&gt;What I care about here is that Local Crew demonstrates a direction I think more AI systems should move toward.&lt;/p&gt;
&lt;p&gt;I want AI infrastructure that is inspectable, composable, and owned by the person using it. I want orchestration that works across the hardware people already have instead of assuming every meaningful workflow has to be rented from somebody else&apos;s platform. I want the system to stay useful when the remote layer is unavailable.&lt;/p&gt;
&lt;p&gt;That&apos;s the reason I keep the emphasis on the local framework. Port is a feature. Auth is a feature. The product is the orchestration environment itself.&lt;/p&gt;
&lt;p&gt;If you want to try it, start locally first. Clone the repo, bootstrap the orchestrator, add a couple of devices, open &lt;code&gt;/ui&lt;/code&gt;, put &lt;code&gt;/display&lt;/code&gt; on a second screen, and run &lt;code&gt;/auto&lt;/code&gt;. Once that local loop is real, Port starts making sense as the remote companion surface.&lt;/p&gt;
&lt;p&gt;That&apos;s the architecture I wanted from the beginning: local-first, operationally legible, and practical enough to run real work on hardware I already control.&lt;/p&gt;</content></entry><entry><title>Cross-Driver Test Architecture: Migrating Enterprise QA with AI Assistance</title><id>https://benlive.tv/blog/cross-driver-test-architecture-migrating-enterprise-qa-with-ai/</id><link href="https://benlive.tv/blog/cross-driver-test-architecture-migrating-enterprise-qa-with-ai/"/><updated>2026-06-29T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Cross-Driver Test Architecture: Migrating Enterprise QA with AI Assistance&lt;/h1&gt;
&lt;p&gt;Most enterprise teams I talk to have the same problem. A large Selenium suite that works, a clear sense that Playwright is the better tool for new development, and no budget to rewrite everything at once. Stopping regression coverage for six months to migrate isn&apos;t an option when the suite is what gates releases.&lt;/p&gt;
&lt;p&gt;At work I&apos;m leading exactly this migration: a VB.NET Selenium codebase moving to C#, on an adapter pattern that implements both a Selenium driver and a Playwright driver behind one interface. I can&apos;t share the code. I can share the architecture and what I&apos;ve learned running it.&lt;/p&gt;
&lt;h2&gt;The adapter is the strategy&lt;/h2&gt;
&lt;p&gt;The core idea is that test logic should never know which browser automation engine is underneath it. A test says &quot;navigate here, find this element, assert this text.&quot; The adapter decides whether that becomes a Selenium &lt;code&gt;WebDriver&lt;/code&gt; call or a Playwright &lt;code&gt;Page&lt;/code&gt; call.&lt;/p&gt;
&lt;p&gt;That gives the team three things at once:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Existing tests keep running on Selenium while they&apos;re ported, so coverage never drops.&lt;/li&gt;&lt;li&gt;New tests are written against the adapter and can run on either driver from day one.&lt;/li&gt;&lt;li&gt;Cross-browser coverage comes from the driver layer, not from rewriting tests.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;The last one matters more than it sounds. On this site&apos;s own suite, three browser projects run the same specs because Playwright abstracts the engine. The adapter brings that property to a codebase that predates Playwright entirely.&lt;/p&gt;
&lt;h2&gt;Two migrations, one track&lt;/h2&gt;
&lt;p&gt;The language move, VB.NET to C#, and the driver move, Selenium toward Playwright, are separate concerns that happen to share a schedule. Keeping them separate in the design is what makes the project tractable. The adapter interface is defined in C#. Ported tests target the interface. Which driver executes them is configuration. A test can be migrated to C# and still run on Selenium the day it lands, which means the port is reviewable on its own without also being a behavior change.&lt;/p&gt;
&lt;h2&gt;Page objects across driver generations&lt;/h2&gt;
&lt;p&gt;The page object model survived the transition, and it&apos;s the reason the migration is possible at all. A page object that exposes intent, &quot;submit the form,&quot; &quot;open the details panel,&quot; instead of exposing locators and waits, ports cleanly because its interface is about the application rather than the driver.&lt;/p&gt;
&lt;p&gt;The page objects that didn&apos;t port cleanly were the ones that had leaked driver calls into test logic over the years. Those got refactored first, and the refactor was the same work whether or not the migration ever happened: it just made the leakage visible.&lt;/p&gt;
&lt;h2&gt;Where assistants help&lt;/h2&gt;
&lt;p&gt;Migration work has a shape assistants are good at: a large volume of similar transformations with a known target pattern. Converting a VB.NET page object to the C# adapter interface is boilerplate once the first one exists. Generating the Playwright implementation of an adapter method when the Selenium one is written is pattern matching. Proposing deterministic waits to replace timing-based ones is a review task with a clear rule.&lt;/p&gt;
&lt;p&gt;So the assistants do that, inside the team&apos;s version-controlled Rules and Profiles, with the adapter&apos;s interface as the contract they can&apos;t change. What they don&apos;t do is decide what the interface is, which tests port first, or when a driver switch is safe. Those decisions are thoroughly human, and they&apos;re the ones that determine whether the migration succeeds.&lt;/p&gt;
&lt;p&gt;The concrete effect has been on maintenance throughput. Script maintenance and enhancement work moves faster when everyone, people and assistants, starts from the same prompt patterns and the same interface.&lt;/p&gt;
&lt;h2&gt;Backward compatibility is non-negotiable&lt;/h2&gt;
&lt;p&gt;The rule I&apos;ve held hardest: at no point during the migration does the release gate get weaker. Every ported test has to pass on the original driver before it&apos;s allowed to run on the new one. Every adapter method has to have parity tests proving both implementations behave the same on the same page. The suite that gates releases today is the suite that gates them tomorrow, plus whatever has been added.&lt;/p&gt;
&lt;p&gt;That rule slows the migration down. It&apos;s also why the team trusts it.&lt;/p&gt;
&lt;h2&gt;Measuring progress&lt;/h2&gt;
&lt;p&gt;Coverage that hasn&apos;t dropped. The count of tests running through the adapter versus directly against Selenium. Defect escape rate holding steady or improving. Execution time, because the Playwright path is faster and the improvement is a visible reward for porting. None of those are vanity numbers. Each one would show a problem early.&lt;/p&gt;
&lt;h2&gt;Why this matters beyond QA&lt;/h2&gt;
&lt;p&gt;Enterprise AI adoption tends to stall on codebases like this one: large, old, and load-bearing. The adapter pattern is what makes such a codebase safe to change at speed, by assistants or by people. Decoupling the volatile layer from the stable one, keeping the gate strong, and letting the boring transformations be automated inside a contract. That&apos;s the same shape as every other piece of AI-assisted engineering I do. This one just happens to be written in C#.&lt;/p&gt;</content></entry><entry><title>Shipping AI Features End to End as a Solo Engineer</title><id>https://benlive.tv/blog/shipping-ai-features-end-to-end-as-a-solo-engineer/</id><link href="https://benlive.tv/blog/shipping-ai-features-end-to-end-as-a-solo-engineer/"/><updated>2026-06-15T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Shipping AI Features End to End as a Solo Engineer&lt;/h1&gt;
&lt;p&gt;The AI Lab on this site has four interactive apps, a serverless backend with model routing and rate limiting, a Playwright suite across three browser engines, a media pipeline that generates social cards from the rendered pages, and a deploy path from commit to production. I built all of it, and I maintain all of it.&lt;/p&gt;
&lt;p&gt;I&apos;m not saying that to brag. I&apos;m saying it because &quot;can ship an AI feature end to end&quot; is a specific skill profile, and the best way to explain it is to walk the loop.&lt;/p&gt;
&lt;h2&gt;Backend&lt;/h2&gt;
&lt;p&gt;Firebase Functions, Node 20, one function per endpoint. The chat function validates the request at the edge, resolves a server-side prompt key, checks the model against a cost allowlist, walks a fallback chain with a circuit breaker, and applies a per-visitor rate limit backed by Firestore. Every AI Lab app talks to the same function with its own prompt keys, so hardening it once hardened all of them.&lt;/p&gt;
&lt;p&gt;The part that separates &quot;it works&quot; from &quot;it ships&quot; is the failure handling. Provider error bodies stay in the logs. Retryable and permanent errors are classified differently at the server and at the client, on purpose. Secrets are Firebase Secrets, and the function refuses to run in production if the IP-hashing salt isn&apos;t bound. None of that is visible to a visitor, and all of it is why the chat has stayed up through provider outages and model retirements.&lt;/p&gt;
&lt;h2&gt;Frontend&lt;/h2&gt;
&lt;p&gt;Vanilla JavaScript, no framework, no build step. That was a deliberate choice for a portfolio: the source a reviewer reads is the source the browser runs. The design system is modern CSS, &lt;code&gt;oklch&lt;/code&gt; color, &lt;code&gt;clamp()&lt;/code&gt; typography, container queries, a light and dark theme applied before first paint, prismatic gradients defined in exactly one file. Accessibility is a requirement: skip links, ARIA patterns for the tier selector and navigation, visible focus rings, reduced-motion handling in the shared animation layer.&lt;/p&gt;
&lt;p&gt;The four apps share modules for configuration, storage, error handling, and transcript rendering, which is where a solo engineer saves the most time. Promptpad, Problem Solver, and Agent Agenda each have their own state model and prompt contracts, and each one inherited the chat&apos;s request contract instead of reinventing it.&lt;/p&gt;
&lt;h2&gt;Testing&lt;/h2&gt;
&lt;p&gt;Fifty-seven spec files. Desktop Chrome, Desktop Firefox, Mobile Safari. A canary of eighteen sentinel tests runs first. A batched runner splits the rest into groups so browsers don&apos;t exhaust the machine. Axe-core scans on twelve pages. Contrast ratios computed in-browser. Pixel-position checks on the navigation across page types. Mocked API responses for the chat so the suite is deterministic. Visual baselines at several viewports.&lt;/p&gt;
&lt;p&gt;The suite is the reason I can change a shared module and know within minutes what it touched. It&apos;s also the thing I trust least on faith, which is why the runner got audited and fixed after I found it reporting killed batches as passing.&lt;/p&gt;
&lt;h2&gt;Media and content&lt;/h2&gt;
&lt;p&gt;The blog renders Markdown client-side, and a script screenshots each rendered article to produce Open Graph, card, and thumbnail images in both themes, recorded in a manifest with signatures so nothing regenerates unnecessarily. Another script produces the sitemap and &lt;code&gt;llms.txt&lt;/code&gt; from one page manifest. Another turns artist-supplied hero PNGs into fifteen responsive variants each. None of this is glamorous. All of it is the difference between a blog and a pile of files.&lt;/p&gt;
&lt;h2&gt;Deploy&lt;/h2&gt;
&lt;p&gt;Firebase Hosting for the static site with strict security headers. Firebase deploy for functions. A predeploy gate that probes each production model with a tiny request. A production smoke script that checks status codes and headers after the deploy. Two git repositories, one public for the site and one private for the infrastructure, with the public one mounted as a directory of the private one.&lt;/p&gt;
&lt;h2&gt;Where assistants change the math&lt;/h2&gt;
&lt;p&gt;The honest answer is that this scope wasn&apos;t reachable for one person on evenings and weekends a few years ago. What changed isn&apos;t that assistants write the code. It&apos;s that they carry context between sessions when the rules are written down, they execute plans I&apos;ve approved while I&apos;m doing something else, and they run the verification loop without getting bored of it.&lt;/p&gt;
&lt;p&gt;The role that&apos;s left for me is the one I&apos;d want anyway: decide what to build, set the constraints, review the plan, review the diff, own the outcome. When that role is done well, one person can hold the whole surface. When it isn&apos;t, no number of assistants helps.&lt;/p&gt;
&lt;h2&gt;What &quot;senior&quot; means here&lt;/h2&gt;
&lt;p&gt;Owning the entire delivery surface changes what senior means. It isn&apos;t knowing the deepest thing about one layer. It&apos;s knowing enough about every layer to make the tradeoffs between them, and being the person accountable when the tradeoff was wrong. The portfolio is the proof I can offer for that: not a list of technologies, but shipped work you can use, with the source open and the tests visible.&lt;/p&gt;</content></entry><entry><title>Production Resilience Patterns for AI-Powered Features</title><id>https://benlive.tv/blog/production-resilience-patterns-for-ai-powered-features/</id><link href="https://benlive.tv/blog/production-resilience-patterns-for-ai-powered-features/"/><updated>2026-06-01T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Production Resilience Patterns for AI-Powered Features&lt;/h1&gt;
&lt;p&gt;The chat on this site looks simple. Pick a tier, send a message, read the answer. Behind that is a resilience system I built because depending on third-party LLM providers means depending on services that return 502s, rate-limit without warning, retire models without notice, and occasionally just stop.&lt;/p&gt;
&lt;p&gt;Everything below is running in production in a single Firebase Function. Each pattern exists because something failed first.&lt;/p&gt;
&lt;h2&gt;The model chain&lt;/h2&gt;
&lt;p&gt;Visitors pick Fast, Balanced, or Thoughtful. Each tier maps to a free-tier model on OpenRouter, and the function holds those in an ordered chain: the Balanced model first, then Thoughtful, then Fast, then OpenRouter&apos;s managed free router as the last resort. If a request fails with a retryable code, the function tries the next model in the chain and logs why.&lt;/p&gt;
&lt;p&gt;The managed router is deliberately last. Pinned models preserve the tier&apos;s character. The router guarantees an answer when the catalog changes underneath me, which it has, with three models retired at once. The chat kept answering. The only visible change was the model name under the replies.&lt;/p&gt;
&lt;h2&gt;The circuit breaker&lt;/h2&gt;
&lt;p&gt;Trying a failing model on every request is a good way to double your latency. The breaker tracks failures per model, opens after three failures within five minutes, and skips open models on later requests until the cooldown passes. The status endpoint exposes the breaker state, so the client can show a tier as unavailable instead of letting a visitor pick it and wait.&lt;/p&gt;
&lt;h2&gt;Two retry policies, and why they differ&lt;/h2&gt;
&lt;p&gt;This one was a real bug. The server has a list of provider error codes that mean &quot;try the next model&quot;: 400, 404, 429, 502, 503, 504. The client has a list of codes worth retrying the whole request: 429, 502, 503, 504.&lt;/p&gt;
&lt;p&gt;For a while the client copied the server&apos;s list verbatim. That meant a 404 from a retired model, which the server had already walked the entire chain to produce, got retried again by the client, consuming another rate-limit slot for a request that couldn&apos;t succeed. The lists are different because the layers are different. By the time the client sees a 400 or 404, every model has already been tried. The rule is now documented in the project&apos;s &lt;code&gt;CLAUDE.md&lt;/code&gt; with the reasoning, so nobody unifies the lists again in the name of tidiness.&lt;/p&gt;
&lt;h2&gt;Cost safety&lt;/h2&gt;
&lt;p&gt;Only allowlisted models run. The allowlist records each model&apos;s provider and the date its free status was verified. A request for anything else throws a cost-safety error before the provider is called. This is the single mechanism that lets me leave a public chat running on a personal account without watching a dashboard.&lt;/p&gt;
&lt;h2&gt;Rate limiting that survives scaling&lt;/h2&gt;
&lt;p&gt;Ten requests per hour per visitor, keyed by an HMAC of the IP under a secret salt so raw addresses are never stored or logged. The window lives in Firestore so every function instance sees the same count, with an in-memory fallback if Firestore is unreachable. The salt is a Firebase Secret and the function refuses to serve in production if it isn&apos;t bound, on the reasoning that hashing with a public default salt over the small IPv4 space isn&apos;t meaningfully different from not hashing at all.&lt;/p&gt;
&lt;p&gt;The limiter is scope-aware now. Device pairing for the Port companion has its own bucket, so a visitor pairing a device never consumes their chat quota, and the two features can&apos;t starve each other.&lt;/p&gt;
&lt;h2&gt;Local parity&lt;/h2&gt;
&lt;p&gt;Every production model has an Ollama stand-in with roughly the same character. When Ollama is detected in development, the function routes to it and the UI shows a local indicator. I test model changes against the stand-in before I deploy them, and the parity table is written into the docs as a requirement rather than a suggestion. Dev and prod that behave differently is how you find out about a retired model from a visitor.&lt;/p&gt;
&lt;h2&gt;What to log&lt;/h2&gt;
&lt;p&gt;Request ID on every response. Which model answered, and whether it was the requested one or a fallback. Circuit breaker transitions. Rate-limit decisions with the hashed key, never the address. Provider error bodies stay in the logs and never reach the browser, because a provider&apos;s error message is provider internals.&lt;/p&gt;
&lt;h2&gt;Degraded, not down&lt;/h2&gt;
&lt;p&gt;The principle under all of it: a visitor asking a question should get an answer from a smaller model before they get an error from a bigger one. Every pattern here serves that. The chain degrades quality before it degrades availability. The breaker protects latency. The allowlist protects cost. The limiter protects the whole thing from one bad actor.&lt;/p&gt;
&lt;p&gt;None of this is exotic. It&apos;s the same resilience engineering any external dependency deserves, applied to a dependency that happens to be a language model.&lt;/p&gt;</content></entry><entry><title>Version-Controlled AI Governance for Engineering Teams</title><id>https://benlive.tv/blog/version-controlled-ai-governance-for-engineering-teams/</id><link href="https://benlive.tv/blog/version-controlled-ai-governance-for-engineering-teams/"/><updated>2026-05-18T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Version-Controlled AI Governance for Engineering Teams&lt;/h1&gt;
&lt;p&gt;Every team using AI assistants has governance, whether or not anyone wrote it down. The only question is whether it&apos;s explicit and reviewable or implicit and different for every engineer.&lt;/p&gt;
&lt;p&gt;At my current job I drafted and iterate on the Rules, Profiles, and Prompts our team uses with Amazon Q Developer, checked in and reviewed like test configuration or deployment scripts. On my own projects the same idea shows up as &lt;code&gt;CLAUDE.md&lt;/code&gt;, &lt;code&gt;AGENTS.md&lt;/code&gt;, and Copilot instruction files. This is what that looks like in practice and why I&apos;d push any team toward it.&lt;/p&gt;
&lt;h2&gt;The governance gap&lt;/h2&gt;
&lt;p&gt;Before the files existed, the team&apos;s AI usage lived in heads. One engineer had a great prompt for triaging a flaky test. Another had learned the hard way not to let the assistant touch the shared fixtures. A third was new and had none of that. The output varied with the person, and the variance showed up in code review as inconsistency nobody could name.&lt;/p&gt;
&lt;p&gt;Writing it down didn&apos;t make anyone&apos;s individual work better. It made everyone&apos;s work the same kind of good.&lt;/p&gt;
&lt;h2&gt;Three kinds of artifact&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Rules&lt;/strong&gt; are engineering constraints the repository owns. Behavior changes require the matching test update. No edits outside the stated task. Deterministic assertions over timing-based waits. Non-trivial releases carry rollback notes. These are reviewed in pull requests. Changing a rule is a team decision with a diff.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Profiles&lt;/strong&gt; describe how the assistant should behave for a kind of work. A regression-triage profile prioritizes finding the root cause over proposing a fix. A script-maintenance profile stays inside the existing patterns of the codebase. Profiles are the reusable behavior, separate from the task at hand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prompts&lt;/strong&gt; are the reusable task templates: diagnose this failure, extend this page object, convert this manual test case. They&apos;re short, they reference the rules, and they exist so nobody rewrites the same instruction from scratch.&lt;/p&gt;
&lt;p&gt;The names come from the tooling we use, but the structure is general. On &lt;a href=&quot;https://github.com/benmcnulty/good-vibes&quot; target=&quot;&lt;em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;good-vibes&lt;/a&gt; it&apos;s three instruction files, one per assistant, each with a stated role. On this site it&apos;s one canonical rule file with a delegation protocol and a set of skills. On &lt;a href=&quot;https://github.com/benmcnulty/startup-cohort&quot; target=&quot;&lt;/em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;startup-cohort&lt;/a&gt;, an experiment in running an AI-agent organization in public, the governance is the whole repo: &lt;code&gt;SOUL.md&lt;/code&gt; for values, &lt;code&gt;AGENTS.md&lt;/code&gt; for execution standards and memory discipline, &lt;code&gt;SECURITY.md&lt;/code&gt; for what may never be committed.&lt;/p&gt;
&lt;h2&gt;Review it like code&lt;/h2&gt;
&lt;p&gt;The change I&apos;d argue for hardest is treating governance edits as code changes. A pull request that loosens a rule gets the same scrutiny as one that loosens an assertion. A new prompt gets read by someone who&apos;ll have to live with its output.&lt;/p&gt;
&lt;p&gt;That&apos;s also where the audit trail comes from. When someone asks why the assistant is allowed to do a thing, the answer is a commit with a reviewer, not a memory of a meeting.&lt;/p&gt;
&lt;h2&gt;How the documents evolve&lt;/h2&gt;
&lt;p&gt;They evolve the same way tests do: from failures. A regression the assistant introduced becomes a rule. A prompt that produced inconsistent output gets tightened. A profile that was too broad gets split. On this site the rule that bugs get a written diagnosis before a fix exists because I watched assistants, and myself, cycle through three CSS approaches without ever naming the property that was wrong.&lt;/p&gt;
&lt;h2&gt;Onboarding is the payoff&lt;/h2&gt;
&lt;p&gt;A new engineer who reads the rules file knows how the team works with AI on day one. So does a new assistant session. That&apos;s the same document doing both jobs, which is the strongest argument I have for keeping it in the repo. It&apos;s also why I built an internal documentation app at work that puts the governance artifacts and the auto-generated codebase docs behind one interface: the onboarding path and the AI context path turned out to be the same path.&lt;/p&gt;
&lt;h2&gt;What changed&lt;/h2&gt;
&lt;p&gt;The measurable effects at work were in the weekly regression cycle. Triage, diagnosis, and resolution got faster and more consistent because everyone started from shared context. Script maintenance improved for the same reason. The less measurable effect was that AI adoption stopped being a set of individual experiments and became a team practice with a definition.&lt;/p&gt;
&lt;p&gt;If I were interviewing an engineer for a team that uses assistants, I&apos;d ask to see their rule files. Not because the files are impressive, but because whether they exist tells you almost everything about how that person works.&lt;/p&gt;</content></entry><entry><title>Onboarding Legacy Codebases for People and AI Assistants</title><id>https://benlive.tv/blog/onboarding-legacy-codebases-for-people-and-ai-assistants/</id><link href="https://benlive.tv/blog/onboarding-legacy-codebases-for-people-and-ai-assistants/"/><updated>2026-05-04T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Onboarding Legacy Codebases for People and AI Assistants&lt;/h1&gt;
&lt;p&gt;At work I built an internal documentation application that consolidates the auto-generated documentation for our codebase into one interface. The goal was to get a new engineer productive on a large, unfamiliar test automation codebase in their first week rather than their first month.&lt;/p&gt;
&lt;p&gt;What I didn&apos;t expect was how much the same tool would do for AI assistants. It turns out that a good onboarding path and good assistant context are almost the same artifact.&lt;/p&gt;
&lt;h2&gt;The problem with legacy onboarding&lt;/h2&gt;
&lt;p&gt;Legacy codebases accumulate knowledge in the wrong places. The reason a helper exists is in a commit message from years ago. The rule about which fixtures are safe to modify is in someone&apos;s head. The auto-generated API docs exist but nobody reads them because they&apos;re a directory of a thousand pages with no path through them.&lt;/p&gt;
&lt;p&gt;New engineers get a tour from whoever has time, take notes, and spend weeks reconstructing the map. Every one of them builds a slightly different map. I can&apos;t publish the codebase, but I can describe the shape of the fix.&lt;/p&gt;
&lt;h2&gt;An app, not more pages&lt;/h2&gt;
&lt;p&gt;The instinct is to write more wiki pages. The problem isn&apos;t a shortage of pages. It&apos;s the absence of a path through them.&lt;/p&gt;
&lt;p&gt;The documentation app puts three things behind one interface: the generated codebase documentation, organized by the modules a new engineer will actually touch first; the team&apos;s AI configuration artifacts, the Rules, Profiles, and Prompts we use daily; and the working guidance for common tasks, like triaging a failed regression run. Search across all of it. A starting point that isn&apos;t &quot;here&apos;s the repo.&quot;&lt;/p&gt;
&lt;p&gt;The effect on onboarding time was measurable. The effect I care more about is consistency: two people onboarded a month apart now hold the same map.&lt;/p&gt;
&lt;h2&gt;The dual audience&lt;/h2&gt;
&lt;p&gt;Here&apos;s the part that changed how I think about documentation.&lt;/p&gt;
&lt;p&gt;When I pointed an assistant at the same consolidated docs, its output on the codebase improved in the same ways a new engineer&apos;s did. It stopped proposing patterns the codebase doesn&apos;t use. It found the right helper instead of writing a new one. It respected the fixture boundaries because they were written down where it would read them.&lt;/p&gt;
&lt;p&gt;That shouldn&apos;t have surprised me. An assistant starting a session is a new engineer with no memory. Everything that helps the human on day one helps the assistant on every day. So the documentation app became, in practice, the team&apos;s context engineering layer, and the two goals stopped being separate projects.&lt;/p&gt;
&lt;h2&gt;What&apos;s worth indexing&lt;/h2&gt;
&lt;p&gt;Not everything. The things that pay off:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;Entry points.&lt;/strong&gt; Where a test run starts, where a page object is registered, where configuration is read. The places a newcomer has to find first.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Boundaries.&lt;/strong&gt; What must not be modified without a conversation. Shared fixtures, environment config, the adapter layer.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Decisions.&lt;/strong&gt; Why the thing is the way it is, so nobody re-argues it. On this site that&apos;s a decision log with numbered records; at work it&apos;s a section of the app.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Task recipes.&lt;/strong&gt; The five things people do every week, as short guides that reference the rules.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;Auto-generated reference docs are the raw material. The index is what makes them usable.&lt;/p&gt;
&lt;h2&gt;The same pattern on my own projects&lt;/h2&gt;
&lt;p&gt;This site has its own version. &lt;code&gt;docs/DOCUMENTATION_MAP.md&lt;/code&gt; is the single index into the rest of the documentation. &lt;code&gt;docs/ONBOARDING.md&lt;/code&gt; is the path for a person. &lt;code&gt;CLAUDE.md&lt;/code&gt; is the path for an assistant, and it points at the same architecture documents. When I reviewed the site recently and found that the two paths had drifted, wrong emulator ports in one, a deleted file still referenced in another, the fix was the same as it would be at work: one index, one source of truth, and a rule that behavior changes update the doc that describes the behavior.&lt;/p&gt;
&lt;p&gt;The legacy corners of this site made the point sharply. The portfolio page is a 2015 jQuery build, and the onboarding for an assistant working on it had to say so explicitly: here is what&apos;s archived, here is what&apos;s current, here is what must stay live for the timeline and what must not be held to the current lint bar. That note is onboarding and context at the same time.&lt;/p&gt;
&lt;h2&gt;The leadership angle&lt;/h2&gt;
&lt;p&gt;Building the documentation app was the point where my role at work shifted from individual contributor toward enabling the team. The lever wasn&apos;t knowing the codebase better than anyone else. It was making that knowledge not depend on me. Gatekeeping expertise looks like job security and functions like a bottleneck. Making the codebase legible to the next person, and the next assistant, is the move that pays off, and it compounds.&lt;/p&gt;
&lt;p&gt;If a codebase can only be understood by talking to a specific person, it can&apos;t be worked on by an assistant either. Fixing one fixes both.&lt;/p&gt;</content></entry><entry><title>Meta Programming the Workflow: Context and Agent Orchestration</title><id>https://benlive.tv/blog/meta-programming-the-workflow-context-and-agent-orchestration/</id><link href="https://benlive.tv/blog/meta-programming-the-workflow-context-and-agent-orchestration/"/><updated>2026-04-21T12:00:00-04:00</updated><content type="html">&lt;h1&gt;Meta Programming the Workflow: Context and Agent Orchestration&lt;/h1&gt;
&lt;p&gt;I spend a lot of time programming products. I spend nearly as much programming the workflow that produces them.&lt;/p&gt;
&lt;p&gt;When a person and several AI assistants work the same branch, quality depends on how context moves between them and how handoffs are structured. Designing that is a real engineering task. I&apos;ve come to think of it as meta programming: writing the rules and feedback loops of the process itself, and checking them in next to the code.&lt;/p&gt;
&lt;h2&gt;What breaks without it&lt;/h2&gt;
&lt;p&gt;Multi-agent velocity without structure produces a specific kind of mess. Duplicated work across contexts that couldn&apos;t see each other. Conflicting edits to shared files. Stale assumptions surviving into a later pass because nothing invalidated them. Merges that are hard to audit because nobody can say which change came from which decision.&lt;/p&gt;
&lt;p&gt;I learned this from the messy merges, not from a blog post. The fix was to design the collaboration layer as deliberately as the product.&lt;/p&gt;
&lt;h2&gt;Roles, written down&lt;/h2&gt;
&lt;p&gt;The pattern that helped most was routing assistants by strength and stating the routing in a file the assistant reads on every session.&lt;/p&gt;
&lt;p&gt;On &lt;a href=&quot;https://github.com/benmcnulty/good-vibes&quot; target=&quot;&lt;em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;good-vibes&lt;/a&gt;, Codex owned planning and the roadmap, Claude Code owned implementation and test coverage, and Copilot reviewed diffs for consistency, each with its own instruction file. On &lt;a href=&quot;https://github.com/benmcnulty/stay-ahead&quot; target=&quot;&lt;/em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;stay-ahead&lt;/a&gt; I pushed the same idea into a test-driven shape: agents produce a plan before code, work through TDD, and keep the documentation synchronized, with the roles defined in the repo. On this site, &lt;code&gt;CLAUDE.md&lt;/code&gt; carries the architecture, the conventions, the test-execution rules, and a delegation protocol: state the technical approach before writing code, diagnose before fixing, size work to fit the session, execute existing plans rather than writing new ones, and run a quality gate before committing.&lt;/p&gt;
&lt;p&gt;Those aren&apos;t aspirations. They&apos;re the failure modes I kept hitting, inverted into rules.&lt;/p&gt;
&lt;h2&gt;Context as a finite resource&lt;/h2&gt;
&lt;p&gt;I treat context the way I&apos;d treat memory in a constrained system. Canonical constraints live in repo docs. Branch-level handoff notes stay current. Validation outcomes get attached to each meaningful change set. Names for files, prompts, tests, and artifacts follow stable conventions so an assistant can find things without being told.&lt;/p&gt;
&lt;p&gt;The payoff shows up when I return to a branch after a few days. The docs say where I left off. I don&apos;t reverse-engineer it from the diff. And a fresh assistant session starts from the same place I do, which is the whole point.&lt;/p&gt;
&lt;h2&gt;Gates with a written pass and fail&lt;/h2&gt;
&lt;p&gt;Review happens in layers, and each layer has an explicit definition of passing: targeted implementation checks, then accessibility and contrast, then integration smoke, then the broader regression run, then a human review before anything deploys. Written criteria make escalation objective. &quot;The canary is green, run the full batch&quot; is a decision anyone can make. &quot;It feels ready&quot; isn&apos;t.&lt;/p&gt;
&lt;h2&gt;Tools I built to run the loop&lt;/h2&gt;
&lt;p&gt;At some point the workflow wanted its own software.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/benmcnulty/hopper&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;hopper&lt;/a&gt; is a visual orchestrator for coding agents. It queues tasks for Claude Code and Codex, runs them in parallel when they&apos;re in different folders and sequentially when they share one, enhances terse task descriptions with a local Ollama model before dispatch, and shows each task&apos;s terminal output in its own card. It exists because I was juggling three terminals and losing track of which one was mid-run.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/benmcnulty/localcrew&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Local Crew&lt;/a&gt; takes the orchestration idea across machines: named agent identities, persistent memory, a task queue filled by consensus, and a verification step before anything is marked complete.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/benmcnulty/orgos&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;orgos&lt;/a&gt; is the most recent and the most ambitious: an agent-first organizational operating system where agent identities, memory, files, and flow sequencing all live on disk as plain JSON and Markdown, so the whole state of the system is readable and diffable.&lt;/p&gt;
&lt;p&gt;Each of those is the same idea at a different scale. Make the process legible. Make the state inspectable. Keep the human deciding what happens next.&lt;/p&gt;
&lt;h2&gt;Automation as coordination&lt;/h2&gt;
&lt;p&gt;Automation isn&apos;t only for speed. It&apos;s how collaborators share a verified view of the system: grouped predeploy test execution with compact reporting, generated JSON summaries of each run, media artifacts with signatures so nothing regenerates unnecessarily, a sitemap and &lt;code&gt;llms.txt&lt;/code&gt; produced from one manifest. Those outputs are the shared state. Anyone, human or assistant, can read them and know where things stand.&lt;/p&gt;
&lt;h2&gt;Documentation is runtime config&lt;/h2&gt;
&lt;p&gt;In a multi-agent environment, a stale document is a bug. Assistants execute against it the same way they&apos;d execute against a wrong constant. So docs update in the same pass as behavior, handoffs carry test evidence, and the decision log records why a thing is the way it is so nobody re-argues it in six months.&lt;/p&gt;
&lt;h2&gt;What I&apos;d tell someone starting this&lt;/h2&gt;
&lt;p&gt;Constrain scope before adding capability. Separate assistant roles by function and write the roles down. Automate verification and artifact generation early, before you need them. Treat context hygiene as a product requirement. Design handoffs so a qualified stranger could continue safely.&lt;/p&gt;
&lt;p&gt;That&apos;s what turns isolated AI productivity wins into throughput you can keep.&lt;/p&gt;</content></entry><entry><title>AI Lab Project Building Part 5: Iterative Engineering with AI Assistants</title><id>https://benlive.tv/blog/ai-lab-project-building-part-5-iterative-engineering-with-ai-assistants/</id><link href="https://benlive.tv/blog/ai-lab-project-building-part-5-iterative-engineering-with-ai-assistants/"/><updated>2026-04-09T12:00:00-04:00</updated><content type="html">&lt;h1&gt;AI Lab Project Building Part 5: Iterative Engineering with AI Assistants&lt;/h1&gt;
&lt;p&gt;By the time the chat, Promptpad, Problem Solver, and Agent Agenda were all running, the most useful thing I&apos;d built wasn&apos;t any one of them. It was the loop that produced them. This post is about that loop, including the parts of it I got wrong and had to fix in a retrospective.&lt;/p&gt;
&lt;h2&gt;The pattern&lt;/h2&gt;
&lt;p&gt;It stayed the same across every app:&lt;/p&gt;
&lt;ol&gt;&lt;li&gt;Write the scope and the non-goals before anything else.&lt;/li&gt;&lt;li&gt;Implement in slices small enough to test on their own.&lt;/li&gt;&lt;li&gt;Run the targeted checks, then the wider ones.&lt;/li&gt;&lt;li&gt;Update the docs and the handoff note in the same pass as the code.&lt;/li&gt;&lt;li&gt;Re-verify with the broader regression run before merge.&lt;/li&gt;&lt;/ol&gt;
&lt;p&gt;Each app got its own feature branch and its own handoff document in &lt;code&gt;docs/&lt;/code&gt;. The handoff listed the files touched, the prompt keys added, the local storage keys, the test layers, and the exact verification commands that had been run. When a branch merged into staging, that document was the reference for what &quot;done&quot; meant.&lt;/p&gt;
&lt;h2&gt;Where the assistants fit&lt;/h2&gt;
&lt;p&gt;I routed by strength rather than treating assistants as interchangeable. Codex handled scoped patch loops with deterministic tests. Claude Code handled architecture interpretation, cross-file risk, and planning. I did scope and final review. That split reduced context thrash, because each tool started from a narrower question, and it made handoffs cleaner, because each produced a different kind of artifact.&lt;/p&gt;
&lt;p&gt;The merges into staging went in clean. Four feature branches, zero conflicts. That wasn&apos;t luck. It was the non-goals list keeping each branch out of the others&apos; files.&lt;/p&gt;
&lt;h2&gt;Frontend: the regressions were small and expensive&lt;/h2&gt;
&lt;p&gt;Almost nothing broke dramatically. What broke was contrast drifting a few points in light mode, a control that looked wired and wasn&apos;t, a layout shifting a pixel after a theme change, focus landing somewhere unexpected after a state transition.&lt;/p&gt;
&lt;p&gt;The practices that cut that down: every visible button is a contract with a test, dark and light parity is verified with a measured contrast ratio rather than a glance, focus management is stateful and keyboard-safe, and typography is tuned for reading rather than for screenshots. Small inconsistencies cost more in perception than their size suggests.&lt;/p&gt;
&lt;p&gt;The About page taught the hardest version of this lesson. Getting a reliable two-page print layout across six resume perspectives took more than fifteen rounds, because &lt;code&gt;page-break&lt;/code&gt; behavior differs between browsers and every round risked regressing the previous one. The thing that made fifteen rounds survivable was locking each component with a regression test before touching the next.&lt;/p&gt;
&lt;h2&gt;Backend: constraints did the work&lt;/h2&gt;
&lt;p&gt;Quality on the server side came from saying no: server-side prompt keys with contracts, model allowlists with a cost-safety check, a controlled fallback chain with structured provider errors, and sequential inference wherever capacity or latency demanded it. No hidden auto-chaining. No speculative background generation. No silent prompt drift. Constrained behavior is easier to reason about and far easier to test.&lt;/p&gt;
&lt;h2&gt;What the retrospective found&lt;/h2&gt;
&lt;p&gt;The February retrospective was honest about the debt. The prismatic gradient that gives the site its look had been added file by file as pages were created, and by the end it was defined six different ways. Navigation parity across four page types with different container widths had eaten significant time. The predeploy pipeline, three parallel tracks, had caught real regressions in the blog reader and the navigation, which validated the investment, but the test runner itself had habits I didn&apos;t discover until later.&lt;/p&gt;
&lt;p&gt;The fix for the gradient was a dedicated sprint: one &lt;code&gt;_gradients.css&lt;/code&gt; as the single source, one &lt;code&gt;_animations.css&lt;/code&gt; for shared keyframes, an import order written into &lt;code&gt;docs/CSS_ARCHITECTURE.md&lt;/code&gt;, and an observer that pauses animations off-screen. Sixty-odd fixes, each with a test. The site looks the same. It&apos;s maintained very differently.&lt;/p&gt;
&lt;h2&gt;Documentation as a runtime dependency&lt;/h2&gt;
&lt;p&gt;The most repeatable lesson: when the docs matched the code, a new session started from decisions. When they were stale, every session began with twenty minutes of rediscovery, and assistants executed against wrong assumptions just as readily as I did.&lt;/p&gt;
&lt;p&gt;So docs became part of the definition of done. Behavior changes update the doc that describes the behavior. Handoffs include test evidence and known risks. The review checklists live in the repo and evolve with it. The decision log records why, not just what, so the next person doesn&apos;t relitigate a settled question.&lt;/p&gt;
&lt;h2&gt;What &quot;done&quot; meant for this phase&lt;/h2&gt;
&lt;p&gt;Each app functionally coherent and covered by layered tests. Regressions constrained by automation rather than memory. Documentation that lets a qualified engineer, or a fresh assistant session, continue safely. A blog that reflects the real state of the work with media generated from the real pages.&lt;/p&gt;
&lt;p&gt;That foundation is what let the next round of changes go faster without going looser.&lt;/p&gt;
&lt;p&gt;Series:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;a href=&quot;/blog/ai-lab-project-building-part-1-ai-chat/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Part 1: AI Chat&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;/blog/ai-lab-project-building-part-2-promptpad/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Part 2: Promptpad&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;/blog/ai-lab-project-building-part-3-problem-solver/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Part 3: Problem Solver&lt;/a&gt;&lt;/li&gt;&lt;li&gt;&lt;a href=&quot;/blog/ai-lab-project-building-part-4-agent-agenda/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Part 4: Agent Agenda&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content></entry><entry><title>AI Lab Project Building Part 4: Agent Agenda</title><id>https://benlive.tv/blog/ai-lab-project-building-part-4-agent-agenda/</id><link href="https://benlive.tv/blog/ai-lab-project-building-part-4-agent-agenda/"/><updated>2026-03-30T12:00:00-04:00</updated><content type="html">&lt;h1&gt;AI Lab Project Building Part 4: Agent Agenda&lt;/h1&gt;
&lt;p&gt;Agent Agenda is the most compositional thing in the AI Lab. The goal is to help you assemble an agent configuration document from reusable blocks rather than writing a wall of instructions from scratch, and to keep you in charge of every block.&lt;/p&gt;
&lt;p&gt;I built it because I was writing these documents by hand for my own repos, the &lt;code&gt;AGENTS.md&lt;/code&gt; and &lt;code&gt;CLAUDE.md&lt;/code&gt; files that tell an assistant what it is and isn&apos;t for, and the structure kept repeating.&lt;/p&gt;
&lt;h2&gt;Blocks, not one big prompt&lt;/h2&gt;
&lt;p&gt;The workbench at &lt;a href=&quot;/ai-lab/agent-workflow/app.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;/ai-lab/agent-workflow/app.html&lt;/a&gt; organizes an agenda into four categories: personas, skills, rules, and tools. Each category holds domains, each domain holds components, and components can link across categories in both directions. A rule can reference the tool it constrains. A skill can reference the persona that owns it.&lt;/p&gt;
&lt;p&gt;Categories can be enabled or disabled per agenda, and the export copies only the enabled ones. That sounds minor until you&apos;re drafting a configuration for two different assistants and want to hand each one only the parts it needs.&lt;/p&gt;
&lt;h2&gt;AI where it helps, and nowhere else&lt;/h2&gt;
&lt;p&gt;There are exactly four generation actions, and every one is a click:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;generate a domain scaffold&lt;/li&gt;&lt;li&gt;draft a single component&lt;/li&gt;&lt;li&gt;fill in default attributes for a component&lt;/li&gt;&lt;li&gt;expand one specific attribute&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;Nothing runs on its own. There&apos;s no &quot;generate my whole agenda&quot; button, on purpose. The backend contract, &lt;code&gt;agentAgenda.generate.v1&lt;/code&gt;, returns strict JSON keyed by action type, so each action produces one predictable shape the UI can merge into the existing document without clobbering what you&apos;ve already written.&lt;/p&gt;
&lt;p&gt;Most agent-builder tools I&apos;ve tried either over-automate or under-structure. Over-automation produces a plausible document you didn&apos;t author and won&apos;t maintain. Under-structure produces a blank page. The middle path is a structured editor with targeted assistance, and that&apos;s what this is.&lt;/p&gt;
&lt;h2&gt;State&lt;/h2&gt;
&lt;p&gt;The agenda persists under &lt;code&gt;ai-lab-agent-agenda-state&lt;/code&gt;. The UI state is explicit and stored with it: which category tab is active, which component is selected, which sections are expanded. Long editing sessions resume where they left off. Manual CRUD works for everything, domains and components alike, so the AI actions are additive rather than required.&lt;/p&gt;
&lt;h2&gt;Testing a composable UI&lt;/h2&gt;
&lt;p&gt;The layering matches the other apps. Unit tests cover parsing, normalization, and the export helper. Integration tests assert the prompt key, the model tier, and the action sequence. The end-to-end spec walks a full workflow, then checks copy and export, accessibility, contrast, and responsive scrolling. A UI-controls spec covers tabs, toggles, manual add and remove, theme, privacy panel, and copy wiring.&lt;/p&gt;
&lt;p&gt;The bidirectional linking got its own attention. Linking component A to component B has to update B&apos;s link list too, and unlinking has to remove both. It&apos;s the kind of invariant that&apos;s easy to get right on the happy path and wrong the moment someone deletes a component that other components point at.&lt;/p&gt;
&lt;h2&gt;Where it points&lt;/h2&gt;
&lt;p&gt;This app is closest to the work I do at my day job, where the team&apos;s AI configuration lives as version-controlled Rules, Profiles, and Prompts. Agent Agenda is the same idea with an editor around it. Structured enough to be consistent across a team, editable enough to reflect real constraints, exportable enough to drop into a repo.&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Previous: &lt;a href=&quot;/blog/ai-lab-project-building-part-3-problem-solver/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;AI Lab Project Building Part 3: Problem Solver&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Next: &lt;a href=&quot;/blog/ai-lab-project-building-part-5-iterative-engineering-with-ai-assistants/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;AI Lab Project Building Part 5: Iterative Engineering with AI Assistants&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content></entry><entry><title>AI Lab Project Building Part 3: Problem Solver</title><id>https://benlive.tv/blog/ai-lab-project-building-part-3-problem-solver/</id><link href="https://benlive.tv/blog/ai-lab-project-building-part-3-problem-solver/"/><updated>2026-03-19T12:00:00-04:00</updated><content type="html">&lt;h1&gt;AI Lab Project Building Part 3: Problem Solver&lt;/h1&gt;
&lt;p&gt;Problem Solver was built around one hard constraint: don&apos;t spend a generation the user didn&apos;t ask for. Every model call is explicit, sequential, and triggered by a click. No background generation, no speculative branching, no &quot;we went ahead and did the next step for you.&quot;&lt;/p&gt;
&lt;p&gt;That constraint made the app slower to use than it could be. It also made it the clearest demonstration in the AI Lab of what I mean when I say human in the loop.&lt;/p&gt;
&lt;h2&gt;The workflow&lt;/h2&gt;
&lt;p&gt;The flow at &lt;a href=&quot;/ai-lab/problem-solving/app.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;/ai-lab/problem-solving/app.html&lt;/a&gt; is deliberately narrow:&lt;/p&gt;
&lt;ol&gt;&lt;li&gt;You describe a rough scenario.&lt;/li&gt;&lt;li&gt;The app proposes a bounded set of candidate problems.&lt;/li&gt;&lt;li&gt;You pick one for analysis.&lt;/li&gt;&lt;li&gt;The app proposes a bounded set of solutions.&lt;/li&gt;&lt;li&gt;You pick one and get an action plan.&lt;/li&gt;&lt;/ol&gt;
&lt;p&gt;Two sliders control how many candidates and how many solutions get generated, one through five, and both default to one. That default alone avoids most of the wasted inference I&apos;ve seen in tools that fan out first and ask later.&lt;/p&gt;
&lt;h2&gt;Four prompt contracts&lt;/h2&gt;
&lt;p&gt;The backend has a separate prompt key for each stage: &lt;code&gt;problemSolver.identify.v1&lt;/code&gt;, &lt;code&gt;problemSolver.analyze.v1&lt;/code&gt;, &lt;code&gt;problemSolver.solutions.v1&lt;/code&gt;, and &lt;code&gt;problemSolver.plan.v1&lt;/code&gt;. Each one is a structured contract with a defined response shape, and each is independently testable. The version suffix is there so I can change a contract without silently changing the behavior of a saved session that was created under the old one.&lt;/p&gt;
&lt;p&gt;Splitting the stages kept every prompt small and specific. One prompt that tries to identify, analyze, and solve in a single pass produces long, hedged output. Four prompts that each do one thing produce output I can assert on.&lt;/p&gt;
&lt;h2&gt;Why sequential gating&lt;/h2&gt;
&lt;p&gt;Three reasons, all practical.&lt;/p&gt;
&lt;p&gt;Both the local and the production inference paths have limited capacity. Sequential steps give me predictable latency and keep token spend proportional to what the visitor actually wanted to see.&lt;/p&gt;
&lt;p&gt;Debugging is cleaner. Each piece of output maps to exactly one request at exactly one step. When something looks wrong, I know which prompt produced it.&lt;/p&gt;
&lt;p&gt;Testing is reproducible. A user path through the app is a fixed sequence of clicks and responses, which means the end-to-end spec can mock each stage and assert the transitions between them.&lt;/p&gt;
&lt;h2&gt;State you can inspect&lt;/h2&gt;
&lt;p&gt;Sessions persist under &lt;code&gt;ai-lab-problem-solver-state&lt;/code&gt;, with create, select, delete, and reset controls. The progression is visible in the UI, and you can return to an earlier validated step without triggering another model call. Copy controls exist for the analysis and the plan, because the output is meant to leave the app.&lt;/p&gt;
&lt;h2&gt;What I tested hardest&lt;/h2&gt;
&lt;p&gt;For a process-heavy app, control state matters at least as much as response quality:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Step transitions are deterministic&lt;/li&gt;&lt;li&gt;Controls stay disabled until their prerequisites are met&lt;/li&gt;&lt;li&gt;Each stage has its own error message when a request fails&lt;/li&gt;&lt;li&gt;No horizontal overflow at the common breakpoints&lt;/li&gt;&lt;li&gt;Labels, control names, and contrast pass in both themes&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;The spec files are layered the same way as the rest of the lab: unit tests for state and parsing helpers, integration tests that assert prompt keys and request order, end-to-end tests for the full workflow, and a UI-controls spec for session management and copy actions.&lt;/p&gt;
&lt;h2&gt;What it proves&lt;/h2&gt;
&lt;p&gt;A constrained inference workflow can still feel capable. Visitors get agency over every branch. I get strong levers on cost and quality. And the app is a working argument for a design position I keep coming back to: the human deciding what happens next isn&apos;t a limitation on the AI, it&apos;s the interface.&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Previous: &lt;a href=&quot;/blog/ai-lab-project-building-part-2-promptpad/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;AI Lab Project Building Part 2: Promptpad&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Next: &lt;a href=&quot;/blog/ai-lab-project-building-part-4-agent-agenda/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;AI Lab Project Building Part 4: Agent Agenda&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content></entry><entry><title>AI Lab Project Building Part 2: Promptpad</title><id>https://benlive.tv/blog/ai-lab-project-building-part-2-promptpad/</id><link href="https://benlive.tv/blog/ai-lab-project-building-part-2-promptpad/"/><updated>2026-03-11T12:00:00-04:00</updated><content type="html">&lt;h1&gt;AI Lab Project Building Part 2: Promptpad&lt;/h1&gt;
&lt;p&gt;Promptpad exists because I kept writing the same terse instruction into a chat box and then spending three turns expanding it. The idea was to treat a prompt like a document: draft it, have a model tighten it, see the diff, keep the version you like.&lt;/p&gt;
&lt;p&gt;There are two Promptpads, and it&apos;s worth being clear about which is which.&lt;/p&gt;
&lt;h2&gt;The standalone app&lt;/h2&gt;
&lt;p&gt;The first one is a Next.js application at &lt;a href=&quot;https://github.com/benmcnulty/promptpad&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;github.com/benmcnulty/promptpad&lt;/a&gt;. It runs entirely against a local Ollama server, so nothing you type leaves your machine. It implements a two-layer refinement pass: a structural rewrite, then an optional normalization pass that only kicks in when a model&apos;s output drifts from the expected shape. It has a CLI for the same workflow, and it carries its own test suite, 156 tests at last count. That app was my first serious attempt at making a local model do one narrow job well, and it taught me that the second pass is where most of the value is. First drafts from small models are fine. Consistency is what they lack, and a normalization step buys a lot of it.&lt;/p&gt;
&lt;h2&gt;The AI Lab version&lt;/h2&gt;
&lt;p&gt;The second one is the workbench at &lt;a href=&quot;/ai-lab/promptpad/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;/ai-lab/promptpad/&lt;/a&gt;. Same idea, different constraints: it runs in a static page against the site&apos;s Firebase Function and OpenRouter, it has to share the chat&apos;s model tiers, and it has to look like it belongs next to the other apps.&lt;/p&gt;
&lt;p&gt;The data model is the part I&apos;d defend most. Every document has an ordered list of revisions. Each revision records its parent, its content, whether it came from a manual edit or an enhancement, which enhancement options were selected, and which model produced it, down to the display name. There&apos;s a committed revision, a history cursor for undo and redo, and a stored &quot;last diff&quot; so the comparison view survives a reload. All of it lives in &lt;code&gt;localStorage&lt;/code&gt; under a versioned key.&lt;/p&gt;
&lt;p&gt;That structure is what makes the workbench feel like a tool instead of a text box. You can enhance, look at what changed, step back, try different options, and commit the one you want.&lt;/p&gt;
&lt;h2&gt;The bug that shaped the sprint&lt;/h2&gt;
&lt;p&gt;The first working version had its own copy of the tier-to-model mapping, forked from the chat&apos;s. The two drifted. Promptpad carried a field the chat didn&apos;t, the chat carried one Promptpad didn&apos;t, and a visitor could pick a tier in Promptpad that sent a request the function would reject. You cannot see a malformed request body in a screenshot, which is exactly how it survived manual review.&lt;/p&gt;
&lt;p&gt;The fix was to delete the fork. The mapping now lives in one shared module, &lt;code&gt;js/api/chat-config.js&lt;/code&gt;, and both apps read it. Promptpad keeps one local table for its own retired-model migrations and one override so its pre-existing storage keys didn&apos;t reset everyone&apos;s saved tier when it moved onto the shared module. The comment above that override explains why it exists, because a year from now it will look like something to clean up.&lt;/p&gt;
&lt;p&gt;Promptpad defaults to the Thoughtful tier where the chat defaults to Balanced. Prompt refinement benefits from the stronger reasoning model more than a quick answer does, and that&apos;s the kind of per-app decision the shared config was built to allow.&lt;/p&gt;
&lt;h2&gt;Wiring every control&lt;/h2&gt;
&lt;p&gt;A second pass went through every visible control and asked whether it did something real. Copy buttons got success feedback. Unlabeled controls were either removed or given an accessible name and a defined behavior. Revision actions had to be deterministic under repeated clicks. The standard was clarity under actual use, not alignment in a mockup.&lt;/p&gt;
&lt;p&gt;Light mode is where visual regressions showed up first, so that&apos;s where I spent the polish: text contrast, panel edges, spacing rhythm matched to the rest of the site.&lt;/p&gt;
&lt;h2&gt;What&apos;s tested&lt;/h2&gt;
&lt;p&gt;Four spec files, layered the same way the other apps are: unit tests for the state serializer and the diff logic, integration tests that assert the prompt key and request order, UI control tests for every button, and an end-to-end run through draft, enhance, diff, commit. If a control is visible, there&apos;s a test that clicks it.&lt;/p&gt;
&lt;h2&gt;What it became&lt;/h2&gt;
&lt;p&gt;Promptpad turned into the style and interaction reference for the two apps that followed. It proved that a sequential, user-triggered inference loop, a real revision model, and a polished editor could ship together and stay testable. That&apos;s harder than it sounds when you&apos;re carrying two themes, several breakpoints, and interactive state across revisions at the same time.&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Previous: &lt;a href=&quot;/blog/ai-lab-project-building-part-1-ai-chat/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;AI Lab Project Building Part 1: AI Chat&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Next: &lt;a href=&quot;/blog/ai-lab-project-building-part-3-problem-solver/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;AI Lab Project Building Part 3: Problem Solver&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content></entry><entry><title>AI Lab Project Building Part 1: AI Chat</title><id>https://benlive.tv/blog/ai-lab-project-building-part-1-ai-chat/</id><link href="https://benlive.tv/blog/ai-lab-project-building-part-1-ai-chat/"/><updated>2026-03-02T12:00:00-05:00</updated><content type="html">&lt;h1&gt;AI Lab Project Building Part 1: AI Chat&lt;/h1&gt;
&lt;p&gt;For a while the chat was the only live thing in the AI Lab, which meant it carried every first impression the site made. That pressure was useful. It forced me to get the unglamorous parts right before adding anything else: how prompts are managed, how models are chosen, how failures are handled, and how all of that is tested.&lt;/p&gt;
&lt;p&gt;Every app that came later, Promptpad, Problem Solver, Agent Agenda, inherited its conventions from this one. So this is where the series starts.&lt;/p&gt;
&lt;h2&gt;What it is&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;/ai-lab/chat/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Chat About Ben&lt;/a&gt; is an interview assistant. It answers questions about my background, projects, and working style, in first person, for people deciding whether to reach out. It&apos;s a static page talking to a Firebase Function, which talks to OpenRouter in production or to Ollama on my machine in development.&lt;/p&gt;
&lt;h2&gt;The rules it established&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The system prompt never leaves the server.&lt;/strong&gt; The client sends &lt;code&gt;promptKey: &quot;aiLab&quot;&lt;/code&gt;, and the function resolves that to the actual prompt. A visitor can&apos;t read the prompt, can&apos;t override it, and can&apos;t send their own system message. Every later app followed the same pattern with its own keys, and the validation is an own-property check now, after I found that a key like &lt;code&gt;&quot;constructor&quot;&lt;/code&gt; would slip past a plain object lookup and quietly drop the prompt.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Only allowlisted models run.&lt;/strong&gt; A short list of verified free-tier models on OpenRouter, each with a verification date. Anything not on the list throws before a request is made. This is the cost-safety mechanism, and it&apos;s why I can leave a public chat running without watching a bill.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Every production model has a local stand-in.&lt;/strong&gt; The three tiers visitors see map to three OpenRouter models, and each of those maps to an Ollama model with roughly the same character. I test locally against the stand-in before I trust the production model. That parity rule is written into the project&apos;s &lt;code&gt;CLAUDE.md&lt;/code&gt; because I&apos;ve broken it by accident and paid for it in surprises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Failures fall through, they don&apos;t surface.&lt;/strong&gt; The function walks a model chain. If the first model returns a retryable error, it tries the next, and the last entry is OpenRouter&apos;s managed free router. A circuit breaker tracks failures per model and skips models that are failing repeatedly. When a model gets retired, which happened in August, the chat keeps answering.&lt;/p&gt;
&lt;h2&gt;The tier selector&lt;/h2&gt;
&lt;p&gt;Visitors get three buttons: Fast, Balanced, and Thoughtful. That&apos;s the entire model-selection UI. Behind it, a stored tier name and a mapping table in &lt;code&gt;functions/model-config.js&lt;/code&gt;. When Ollama is detected in development, a fourth Custom tier appears with whatever local models are installed.&lt;/p&gt;
&lt;p&gt;The selector had to handle its own history. Earlier versions stored a raw model ID. When those IDs were retired, the migration mapped old IDs onto tiers so nobody&apos;s saved preference turned into an error. Small thing, easy to skip, and it&apos;s exactly the kind of detail that separates a demo from a product.&lt;/p&gt;
&lt;p&gt;One honest note: when the Balanced model is unavailable and the chain falls through, the response badge shows the model that actually answered. I&apos;d rather the visitor see &quot;Nemotron&quot; under a reply than believe Gemma wrote it.&lt;/p&gt;
&lt;h2&gt;Where my time went&lt;/h2&gt;
&lt;p&gt;Not on the model. On the surrounding behavior:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Rate limiting per IP, with the IP hashed under a secret salt before it&apos;s stored, and a shared Firestore-backed window so multiple function instances agree&lt;/li&gt;&lt;li&gt;Retry logic on the client for the status codes that are worth retrying, and only those, because the server has already walked the whole chain by the time the client sees a 400&lt;/li&gt;&lt;li&gt;A markdown renderer that escapes rather than executes, after a review flagged &lt;code&gt;innerHTML&lt;/code&gt; as a risk surface&lt;/li&gt;&lt;li&gt;Visible loading states, a retry path, and error messages that say something useful&lt;/li&gt;&lt;/ul&gt;
&lt;h2&gt;How it&apos;s tested&lt;/h2&gt;
&lt;p&gt;The chat UI spec alone is sixty tests, all against a mocked API so they&apos;re deterministic: rendering, markdown, the tier selector&apos;s ARIA radiogroup pattern, keyboard focus after selecting a tier, contrast in both themes, mobile viewport behavior. Separate specs cover the shared modules, the rate limiter&apos;s behavior under unique IPs, and the request validation.&lt;/p&gt;
&lt;p&gt;The rule that built that suite was simple. Every bug I found by hand got an automated check so it couldn&apos;t come back. Over time that produces a suite shaped like how people actually break things, rather than how I imagined they might.&lt;/p&gt;
&lt;h2&gt;Why start here&lt;/h2&gt;
&lt;p&gt;Promptpad inherited the tier selector and the request contract. Problem Solver and Agent Agenda inherited the explicit, user-triggered inference pattern and the test layering. Getting the first one steady made every one after it cheaper.&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Next: &lt;a href=&quot;/blog/ai-lab-project-building-part-2-promptpad/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;AI Lab Project Building Part 2: Promptpad&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;</content></entry><entry><title>Automating Quality Assurance for AI-Powered Development</title><id>https://benlive.tv/blog/automating-quality-assurance-for-ai-powered-development/</id><link href="https://benlive.tv/blog/automating-quality-assurance-for-ai-powered-development/"/><updated>2026-02-23T12:00:00-05:00</updated><content type="html">&lt;h1&gt;Automating Quality Assurance for AI-Powered Development&lt;/h1&gt;
&lt;p&gt;AI assistants write code fast. That&apos;s the easy part. The hard part is making sure their output actually ships: it doesn&apos;t break existing behavior, doesn&apos;t drift from project conventions, and doesn&apos;t quietly introduce regressions that surface three deploys later.&lt;/p&gt;
&lt;p&gt;I&apos;ve spent the past several months building a QA workflow that treats AI-assisted code with the same rigor I&apos;d apply to any contributor&apos;s work. This article walks through the specific tools, patterns, and automation I use to keep quality high when multiple AI agents are touching the same codebase.&lt;/p&gt;
&lt;h2&gt;Engineering Discipline as Project Configuration&lt;/h2&gt;
&lt;p&gt;The biggest productivity lever isn&apos;t how you talk to an AI assistant in the moment. It&apos;s what the assistant already knows before the conversation starts.&lt;/p&gt;
&lt;p&gt;I maintain a layered set of project rules that every AI session inherits automatically. At the global level, a delegation protocol defines how work gets structured: state the specific technical approach before writing code (not just the desired outcome), produce a written diagnosis before fixing any bug, size each task to produce a shippable increment within the session, and run a quality gate before committing. At the project level, a CLAUDE.md file encodes architectural decisions, CSS conventions, test execution rules, and the dual-repo workflow.&lt;/p&gt;
&lt;p&gt;I treat this as configuration rather than ad hoc prompting. The rules are checked into the project alongside the code they govern, and they persist across sessions, tools, and contributors, human or AI. When I start a new Claude Code session on a branch I haven&apos;t touched in a week, the assistant already knows that &lt;code&gt;bun&lt;/code&gt; is the runtime, that Playwright tests run in batches rather than one bulk run, that CSS gradients live in &lt;code&gt;_gradients.css&lt;/code&gt; and nowhere else, and that &lt;code&gt;transition: all&lt;/code&gt; is banned.&lt;/p&gt;
&lt;p&gt;The practical difference is significant. Instead of restating project conventions in every conversation, I invest that effort once into the configuration files and then benefit from it indefinitely. When the rules need to change (say we adopt a new test naming convention), I update the configuration, not my prompting habits.&lt;/p&gt;
&lt;p&gt;A few principles guide how I write these rules:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;I write spec constraints, not outcomes. &quot;Use &lt;code&gt;margin-left: auto&lt;/code&gt; on &lt;code&gt;.nav&lt;/code&gt; inside a flex parent&quot; is enforceable. &quot;Right-align the nav&quot; can produce different implementations across sessions.&lt;/li&gt;&lt;li&gt;For bugs, diagnosis comes first. The rule requires a ranked hypothesis list before any code changes, which reduces wrong-root-cause fixes.&lt;/li&gt;&lt;li&gt;Evidence beats adjectives. &quot;WCAG AA requires 4.5:1 contrast ratio&quot; is verifiable. &quot;Make it more readable&quot; is not.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;I also keep exit criteria explicit: &quot;Done when the Playwright accessibility spec passes on Desktop Chrome and Mobile Safari&quot; removes ambiguity about when to stop iterating.&lt;/p&gt;
&lt;p&gt;The result is that most of my interactions with AI assistants are short and surgical. The long conversations are the ones where I&apos;m updating the rules themselves, and those pay dividends across every future session.&lt;/p&gt;
&lt;h2&gt;Version Control as a Safety Net&lt;/h2&gt;
&lt;p&gt;When AI assistants are generating multi-file changes, version control discipline becomes non-negotiable. I commit in small, scoped batches (one logical change per commit) with conventional commit messages that make &lt;code&gt;git log --oneline&lt;/code&gt; useful six months later.&lt;/p&gt;
&lt;p&gt;The dual-repo structure of this project (root repo for functions, tests, and config; &lt;code&gt;public/&lt;/code&gt; as its own repo for static assets) forces me to be explicit about what changed where. That friction is actually helpful. It prevents the &quot;commit everything and sort it out later&quot; pattern that makes rollbacks painful.&lt;/p&gt;
&lt;p&gt;Branch hygiene matters more with AI-assisted development, not less. When an assistant refactors a CSS file, I want that isolated on a feature branch where I can review the diff, run targeted tests, and merge with confidence, not mixed into an unrelated feature commit.&lt;/p&gt;
&lt;h2&gt;Test Coverage That Actually Catches Things&lt;/h2&gt;
&lt;p&gt;This project currently runs 46 Playwright spec files across three browser projects: Desktop Chrome, Desktop Firefox, and Mobile Safari. That sounds like a lot of tests, and it is. But the quantity isn&apos;t the point. The structure is.&lt;/p&gt;
&lt;p&gt;I organize tests by concern, not by page. There are specs for navigation consistency across every page, CSS regression contracts for gradient tokens and animation timing, accessibility scans using axe-core on 12 different pages, responsive scaling assertions for breakpoints from 320px through 4K, and API endpoint validation for the chat backend.&lt;/p&gt;
&lt;p&gt;The test that catches the most regressions isn&apos;t the biggest or most sophisticated. It&apos;s the one that checks whether every page&apos;s navigation renders at the same pixel position. When an assistant changes a padding value in one CSS file, that test lights up immediately.&lt;/p&gt;
&lt;h2&gt;Canary Tests: Fast Feedback Before the Full Suite&lt;/h2&gt;
&lt;p&gt;Running 46 spec files across 3 browser projects takes time. When I&apos;m iterating quickly with an AI assistant, I don&apos;t want to wait for the full suite after every change.&lt;/p&gt;
&lt;p&gt;That&apos;s where canary tests come in. I have 18 sentinel tests in a single spec file that sample every area of the site: landing page rendering, blog card counts, AI Lab hub layout, chat UI controls, navigation alignment, accessibility basics, API health, and SEO asset presence.&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;bash&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;bun run test:canary&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;If all 18 pass, I haven&apos;t broken anything obvious and can keep moving. If one fails, I know exactly which area to investigate before running the full suite. The canary run completes in a fraction of the time, which means I actually run it. That&apos;s the real value.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/automating-quality-assurance-for-ai-powered-development/illustration-1-dark-960w.jpg&quot; alt=&quot;Geometric illustration of a canary sentinel standing watch over a grid of 18 passing test indicators&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;h2&gt;Batched Test Execution: Dot-and-F Output&lt;/h2&gt;
&lt;p&gt;When I do need the full suite, I don&apos;t run all 46 spec files in a single Playwright invocation. That approach creates resource pressure, orphaned browser processes, and output so long it&apos;s useless for quick assessment.&lt;/p&gt;
&lt;p&gt;Instead, I use a batched runner that groups spec files into 14 sequential batches by area:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;bash&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;bun run test:batched&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Each batch runs with Playwright&apos;s dot reporter and produces a single compact summary line:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;text&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;canary             ..................  ok   (12s)
landing+hire-me    ...............     ok   (18s)
ai-lab-core        ..........FF..     FAIL (25s)&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The dot-and-F format gives me at-a-glance pass/fail across the entire suite without scrolling through verbose output. A dot means pass, an F means failure, and I can immediately see which batch needs attention.&lt;/p&gt;
&lt;p&gt;There&apos;s also a canary-gated mode (&lt;code&gt;bun run test:batched:canary&lt;/code&gt;) that runs the canary batch first and exits early if everything passes, useful when I&apos;m confident a change is low-risk and want fast confirmation.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/automating-quality-assurance-for-ai-powered-development/illustration-2-dark-960w.jpg&quot; alt=&quot;Top-down view of 9 sequential test batches flowing through a pipeline with dot-and-F pass/fail notation&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;h2&gt;Claude Code: Skills, Rules, and Institutional Memory&lt;/h2&gt;
&lt;p&gt;Claude Code&apos;s value for QA goes deeper than code generation. It&apos;s the system I use to encode how this project works: what to build, how to build it, how to test it, and when to stop and think before acting.&lt;/p&gt;
&lt;p&gt;The skill system turns recurring workflows into reusable commands. I have ten skills covering the development lifecycle: build, dev server, testing, dependency management, formatting, type checking, cleanup, &lt;strong&gt;diagnostics&lt;/strong&gt;, &lt;strong&gt;audit&lt;/strong&gt;, and &lt;strong&gt;plan execution&lt;/strong&gt;. The last three are the ones that matter most for quality.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;/diagnose&lt;/code&gt; skill enforces hypothesis-driven debugging. When something breaks, the skill requires a written ranked hypothesis list, three candidates from most to least likely, before any code changes happen. This prevents the most expensive AI failure mode: confidently applying a fix to the wrong root cause, then spending the next hour unwinding cascading side effects.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;/audit&lt;/code&gt; skill is the quality gate I run before committing significant changes. It reviews recent changes against project standards: CSS architecture compliance, accessibility patterns, security considerations, test coverage. Skipping it is the exception, not the norm.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;/implement-plan&lt;/code&gt; skill picks up an existing plan document and executes it phase by phase. In that workflow, each phase is verified and committed before moving on. This prevents another common failure mode: an AI assistant that re-analyzes the entire codebase every time you resume work on a feature, producing a new plan that subtly diverges from the one you already approved.&lt;/p&gt;
&lt;p&gt;But skills are just the command layer. The real institutional memory lives in the project&apos;s CLAUDE.md file: a structured document that encodes architectural decisions, CSS conventions, test execution rules, the dual-repo workflow, API resilience patterns, and the delegation protocol itself. When I start a new session on any branch, the assistant inherits all of this context automatically. No re-explaining that &lt;code&gt;bun&lt;/code&gt; is the runtime, no reminding it about the CSS import chain, no restating that tests run in batches.&lt;/p&gt;
&lt;p&gt;Used together, skills and project rules create the same kind of consistency teams usually enforce through onboarding docs and code review, but with less session-to-session drift.&lt;/p&gt;
&lt;h2&gt;GitHub Copilot in VS Code: IDE-Native AI for This Project&lt;/h2&gt;
&lt;p&gt;GitHub Copilot is where most of the IDE-integrated AI work on this project happens. In Visual Studio Code, I typically use a Claude-family model for planning and heavier reasoning, then switch to Codex-family models for fast implementation passes. That split maps well to real tasks: architecture and dependency analysis first, then rapid edits on known patterns.&lt;/p&gt;
&lt;p&gt;The practical QA value is the tight loop inside the editor. I can move from failing test output to file search to patching in one place without context switching. For responsive layout work where I&apos;m checking computed styles, comparing breakpoint behavior, and validating accessibility attributes, that density matters.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/automating-quality-assurance-for-ai-powered-development/illustration-3-dark-960w.jpg&quot; alt=&quot;Three AI agents, Claude Code, GitHub Copilot, and the developer, collaborating inside a unified development workspace&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;h2&gt;Codex as an Equal Development Partner&lt;/h2&gt;
&lt;p&gt;Codex isn&apos;t an assistant I occasionally delegate to. It&apos;s a co-developer that ships production code on this project daily. It handles the staging audit that caught four test issues across 46 spec files, applies responsive CSS fixes that pass the full suite on the first run, and commits with conventional messages scoped to the right repo. When the conventions are encoded in &lt;code&gt;AGENTS.md&lt;/code&gt;, Codex operates inside them reliably.&lt;/p&gt;
&lt;p&gt;What makes this partnership productive is how naturally Codex handles the full implementation cycle. A typical handoff isn&apos;t &quot;fix this bug.&quot; It&apos;s &quot;diagnose the rate-limit test flake, verify the fix with curl against both emulator ports, update the test, run the canary suite, then commit.&quot; Codex executes that entire sequence autonomously: it reads the function source to understand IP hashing, runs diagnostic curl loops, identifies the root cause (PID-only seeding causing cross-run collisions), applies the fix, and verifies it end-to-end before committing. That&apos;s collaboration through delegation to a developer who excels in a terminal.&lt;/p&gt;
&lt;p&gt;The planning-to-implementation handoff is smooth when the plan names exact files, commands, and acceptance checks. Friction shows up when plan language is broad (&quot;tighten this section&quot; or &quot;make this cleaner&quot;) because the model has to infer tone and boundaries that were never explicit. The pattern that works is detailed planning, narrow implementation scopes, and frequent verification: the same discipline that makes any engineering partnership productive.&lt;/p&gt;
&lt;h2&gt;Amazon Q Developer: Enterprise QA at Scale&lt;/h2&gt;
&lt;p&gt;Amazon Q Developer occupies a different role in my workflow. I use it in my enterprise QA automation work in Visual Studio Professional, where the context is government-scale test infrastructure rather than a personal project. The specifics of that work are proprietary, but the patterns are worth discussing at a general level.&lt;/p&gt;
&lt;p&gt;Q Developer&apos;s value in an enterprise setting is in its governance layer. Rules, profiles, and organizational customizations let teams define coding standards that are enforced during code generation, not discovered during review. For large QA automation codebases where consistency across contributors matters as much as correctness, that enforcement-at-generation approach reduces the review burden significantly.&lt;/p&gt;
&lt;p&gt;Q&apos;s customization and tooling organization pattern differs from Copilot&apos;s model. It&apos;s designed around enterprise team structures rather than individual developer workflows. The profile system supports switching between project contexts cleanly, and the rules framework can encode team-specific conventions that persist across sessions and contributors.&lt;/p&gt;
&lt;p&gt;I mention Q here not because it touches this project&apos;s codebase (it doesn&apos;t) but because the discipline of working with enterprise-grade AI governance tools informs how I think about quality everywhere. The habit of encoding standards as machine-readable rules rather than tribal knowledge is transferable regardless of which IDE or assistant you&apos;re using.&lt;/p&gt;
&lt;h2&gt;The Architect and the Builder: Plan-Then-Implement&lt;/h2&gt;
&lt;p&gt;The single biggest QA improvement in my workflow is separating planning from implementation as distinct operational modes with different capabilities.&lt;/p&gt;
&lt;p&gt;Claude Code&apos;s &lt;code&gt;/model opusplan&lt;/code&gt; configuration enables a structured plan-then-implement pattern. In my workflow, planning is a non-mutating phase: explore the codebase, search for patterns, read files, and reason through dependencies before touching implementation files. The output is a plan document: a concrete, reviewable artifact that specifies which files will change, what approach will be taken, what risks exist, and how to verify the result.&lt;/p&gt;
&lt;p&gt;That plan document is the QA checkpoint. Before a single line of code changes, I can read the plan and catch the problems that cause regressions: the CSS cascade effect of changing a shared variable, the test that depends on a specific DOM structure, the other page that imports the same component. The plan makes those dependencies visible before any files are touched.&lt;/p&gt;
&lt;p&gt;Once I approve the plan, implementation begins, often with a different model optimized for fast, focused code generation. The implementing model receives the plan as its specification and executes it step by step, following the &lt;code&gt;/implement-plan&lt;/code&gt; workflow of phase-level verification and commits.&lt;/p&gt;
&lt;p&gt;This separation works because planning and implementation require different strengths. Planning needs broad codebase awareness, careful reasoning about side effects, and the patience to explore before proposing. Implementation needs speed, precision, and the discipline to follow a spec without scope creep. Using the right tool for each phase produces better results than asking one model to do both simultaneously.&lt;/p&gt;
&lt;p&gt;The extended thinking feature makes the planning phase transparent: I can see which files the model considered, what tradeoffs it identified, and why it chose one approach over another. That visibility turns plan review from a rubber stamp into a genuine quality gate. I&apos;ve caught architectural mistakes in plans that would have cost hours to unwind if they&apos;d gone straight to implementation.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/automating-quality-assurance-for-ai-powered-development/illustration-4-dark-960w.jpg&quot; alt=&quot;Split composition showing planning phase on the left, blueprints and decision trees, flowing into implementation phase on the right: code, green test dots, and deploy&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;h2&gt;Putting It All Together&lt;/h2&gt;
&lt;p&gt;None of these tools work in isolation. The QA workflow that actually keeps this project shippable is the combination:&lt;/p&gt;
&lt;ol&gt;&lt;li&gt;&lt;strong&gt;Encode discipline as configuration&lt;/strong&gt;: CLAUDE.md rules, delegation protocol, project conventions&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Plan before implementing&lt;/strong&gt;: non-mutating exploration, reviewable plan document, then focused execution&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Commit in small scopes&lt;/strong&gt;: one logical change, conventional message&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Run canary after every change&lt;/strong&gt;: 18 tests, fast feedback&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Run full batched suite before merge&lt;/strong&gt;: 14 batches, dot-and-F assessment&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Use IDE integration where it fits&lt;/strong&gt;: Copilot for tight feedback loops, Q for enterprise governance&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Enforce quality gates with skills&lt;/strong&gt;: &lt;code&gt;/diagnose&lt;/code&gt; before fixing, &lt;code&gt;/audit&lt;/code&gt; before committing, &lt;code&gt;/implement-plan&lt;/code&gt; for plan execution&lt;/li&gt;&lt;/ol&gt;
&lt;p&gt;The insight that took me longest to learn: AI-assisted development doesn&apos;t need less QA. It needs more, but it also gives you better tools to automate it. The same technology that generates code quickly can also validate it quickly, if you invest in the infrastructure.&lt;/p&gt;
&lt;p&gt;Quality isn&apos;t something you add at the end. It&apos;s the workflow itself.&lt;/p&gt;</content></entry><entry><title>Building a Dual-Source Blog Engine with Responsive Media</title><id>https://benlive.tv/blog/building-a-dual-source-blog-engine-with-responsive-media/</id><link href="https://benlive.tv/blog/building-a-dual-source-blog-engine-with-responsive-media/"/><updated>2026-02-23T12:00:00-05:00</updated><content type="html">&lt;style&gt;
  /* Article-specific enhancements: layered on top of blog.css defaults */
  .pipeline-diagram {
    display: grid;
    grid-template-columns: 1fr auto 1fr;
    align-items: center;
    gap: 1.5rem;
    margin: 2.5rem 0;
    padding: 2rem;
    border-radius: var(--radius-lg, 0.75rem);
    background: rgba(15, 23, 42, 0.4);
    border: 1px solid rgba(148, 163, 184, 0.15);
  }

  html[data-theme=&quot;light&quot;] .pipeline-diagram {
    background: rgba(248, 250, 252, 0.8);
    border-color: rgba(15, 23, 42, 0.1);
  }

  .pipeline-source,
  .pipeline-output {
    display: flex;
    flex-direction: column;
    gap: 0.75rem;
  }

  .pipeline-arrow {
    font-size: 2rem;
    color: var(--accent-cyan, #22d3ee);
    text-align: center;
    line-height: 1;
  }

  .format-badge {
    display: inline-flex;
    align-items: center;
    gap: 0.4rem;
    padding: 0.35rem 0.75rem;
    border-radius: 0.375rem;
    font-size: 0.82rem;
    font-weight: 600;
    font-family: var(--font-mono, &apos;SF Mono&apos;, monospace);
    letter-spacing: 0.02em;
  }

  .format-badge.md {
    background: rgba(34, 211, 238, 0.15);
    color: var(--accent-cyan, #22d3ee);
    border: 1px solid rgba(34, 211, 238, 0.3);
  }

  .format-badge.html {
    background: rgba(124, 58, 237, 0.15);
    color: var(--accent-purple, #7c3aed);
    border: 1px solid rgba(124, 58, 237, 0.3);
  }

  .format-badge.webp {
    background: rgba(52, 211, 153, 0.15);
    color: var(--accent-green, #34d399);
    border: 1px solid rgba(52, 211, 153, 0.3);
  }

  .format-badge.jpg {
    background: rgba(251, 191, 36, 0.15);
    color: var(--accent-yellow, #fbbf24);
    border: 1px solid rgba(251, 191, 36, 0.3);
  }

  .architecture-table {
    width: 100%;
    border-collapse: collapse;
    margin: 1.5rem 0;
    font-size: 0.92rem;
  }

  .architecture-table th,
  .architecture-table td {
    padding: 0.65rem 0.85rem;
    text-align: left;
    border-bottom: 1px solid rgba(148, 163, 184, 0.15);
  }

  .architecture-table th {
    font-weight: 600;
    color: var(--text-primary, #f1f5f9);
    background: rgba(15, 23, 42, 0.3);
  }

  html[data-theme=&quot;light&quot;] .architecture-table th {
    color: var(--text-primary, #0f172a);
    background: rgba(241, 245, 249, 0.8);
  }

  .architecture-table tr:hover td {
    background: rgba(148, 163, 184, 0.06);
  }

  .code-comparison {
    display: grid;
    grid-template-columns: 1fr 1fr;
    gap: 1rem;
    margin: 1.5rem 0;
  }

  .code-comparison &gt; div {
    overflow: hidden;
    border-radius: var(--radius-md, 0.5rem);
  }

  .code-comparison pre {
    margin: 0;
    padding: 1rem;
    font-size: 0.82rem;
    line-height: 1.55;
    overflow-x: auto;
    background: rgba(10, 15, 26, 0.6);
    border: 1px solid rgba(148, 163, 184, 0.12);
    border-radius: var(--radius-md, 0.5rem);
  }

  html[data-theme=&quot;light&quot;] .code-comparison pre {
    background: rgba(241, 245, 249, 0.9);
    border-color: rgba(15, 23, 42, 0.1);
  }

  .code-comparison .label {
    display: block;
    padding: 0.4rem 0.75rem;
    font-size: 0.75rem;
    font-weight: 600;
    text-transform: uppercase;
    letter-spacing: 0.06em;
    color: var(--text-muted, #94a3b8);
  }

  .width-variants {
    display: flex;
    gap: 0.5rem;
    flex-wrap: wrap;
    margin: 0.5rem 0;
  }

  .width-pill {
    padding: 0.25rem 0.6rem;
    border-radius: 1rem;
    font-size: 0.78rem;
    font-weight: 500;
    font-family: var(--font-mono, &apos;SF Mono&apos;, monospace);
    background: rgba(148, 163, 184, 0.12);
    color: var(--text-secondary, #cbd5e1);
    border: 1px solid rgba(148, 163, 184, 0.2);
  }

  @media (max-width: 640px) {
    .pipeline-diagram {
      grid-template-columns: 1fr;
      text-align: center;
    }

    .pipeline-arrow {
      transform: rotate(90deg);
    }

    .code-comparison {
      grid-template-columns: 1fr;
    }
  }
&lt;/style&gt;

&lt;h1&gt;Building a Dual-Source Blog Engine with Responsive Media&lt;/h1&gt;

&lt;p&gt;Most developer blogs pick a format and stick with it. Markdown is the default: write in plain text, render to HTML, ship. It works. But there are articles where Markdown&apos;s simplicity becomes a constraint: interactive demos, custom layouts, embedded visualizations, or any content where you need precise control over the DOM.&lt;/p&gt;

&lt;p&gt;This blog now supports both. Markdown articles and HTML articles flow through the same content pipeline, share the same frontmatter system, and render inside the same page shell with all the site&apos;s default styling intact. An HTML article only needs to provide its custom overrides; everything else comes for free.&lt;/p&gt;

&lt;p&gt;This post explains how that works, and also covers the responsive media pipeline that processes artist-generated illustrations into pixel-perfect variants for every screen size.&lt;/p&gt;

&lt;h2&gt;The Content Architecture&lt;/h2&gt;

&lt;p&gt;Every article on this blog is a file in &lt;code&gt;/content/blog/&lt;/code&gt; registered in a central &lt;code&gt;index.json&lt;/code&gt; manifest. The content loader reads each file, extracts frontmatter, and either parses the body as Markdown or injects it directly as HTML, decided entirely by file extension.&lt;/p&gt;

&lt;div class=&quot;pipeline-diagram&quot;&gt;
  &lt;div class=&quot;pipeline-source&quot;&gt;
    &lt;span class=&quot;format-badge md&quot;&gt;.md&lt;/span&gt;
    &lt;span class=&quot;format-badge html&quot;&gt;.html&lt;/span&gt;
  &lt;/div&gt;
  &lt;div class=&quot;pipeline-arrow&quot;&gt;→&lt;/div&gt;
  &lt;div class=&quot;pipeline-output&quot;&gt;
    &lt;span&gt;Frontmatter extraction&lt;/span&gt;
    &lt;span&gt;Format-specific rendering&lt;/span&gt;
    &lt;span&gt;Unified article shell&lt;/span&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;p&gt;The key design decision: both formats use the same &lt;code&gt;---&lt;/code&gt; frontmatter block. This means the listing page, SEO meta tags, social sharing cards, and media resolution all work identically regardless of whether the article body is Markdown or HTML.&lt;/p&gt;

&lt;h2&gt;Markdown: The Default Path&lt;/h2&gt;

&lt;p&gt;Most articles are Markdown. The content loader fetches the file, splits the frontmatter, and passes the body through a custom parser that converts it to semantic HTML. The parser handles headings, code blocks with language classes and copy buttons, tables, blockquotes, lists, inline formatting, images, and links.&lt;/p&gt;

&lt;p&gt;The parser also escapes all user content before wrapping it in HTML tags, a security measure that prevents any accidental script injection from Markdown source files. This is the right tradeoff for Markdown: the format is meant to be simple and safe.&lt;/p&gt;

&lt;h2&gt;HTML: Full Control When You Need It&lt;/h2&gt;

&lt;p&gt;HTML articles bypass the Markdown parser entirely. After frontmatter extraction, the body is injected directly into the article container via &lt;code&gt;innerHTML&lt;/code&gt;. This means you get the full power of HTML and CSS inside the article: inline &lt;code&gt;&amp;lt;style&amp;gt;&lt;/code&gt; tags for custom styling and any HTML structure you need. Note that &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt; tags are preserved in the DOM but not executed, since &lt;code&gt;innerHTML&lt;/code&gt; does not evaluate injected scripts per the HTML specification.&lt;/p&gt;

&lt;p&gt;The article you&apos;re reading now is an HTML article. That pipeline diagram above? It&apos;s a CSS Grid layout with custom-styled badges, something that would be clumsy to express in Markdown. The code comparisons below use a side-by-side grid that would be impossible in plain Markdown.&lt;/p&gt;

&lt;h3&gt;What You Get for Free&lt;/h3&gt;

&lt;p&gt;HTML articles inherit the full blog styling system automatically:&lt;/p&gt;

&lt;table class=&quot;architecture-table&quot;&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Feature&lt;/th&gt;
      &lt;th&gt;Source&lt;/th&gt;
      &lt;th&gt;Override Needed?&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Typography (headings, body, code)&lt;/td&gt;
      &lt;td&gt;blog.css&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Dark/light theme switching&lt;/td&gt;
      &lt;td&gt;theme-toggle.js + CSS variables&lt;/td&gt;
      &lt;td&gt;Only for custom elements&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Responsive layout + reader scaling&lt;/td&gt;
      &lt;td&gt;blog.css + article.js&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Hero image with srcset + WebP&lt;/td&gt;
      &lt;td&gt;article.js + frontmatter&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SEO meta + Open Graph + JSON-LD&lt;/td&gt;
      &lt;td&gt;article.js + blog-utils.js&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Reading progress + time estimate&lt;/td&gt;
      &lt;td&gt;article.js&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Heading permalinks + copy links&lt;/td&gt;
      &lt;td&gt;article.js&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Theme-aware inline images&lt;/td&gt;
      &lt;td&gt;article.js&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;An HTML article&apos;s inline &lt;code&gt;&amp;lt;style&amp;gt;&lt;/code&gt; block only needs rules for &lt;em&gt;its own custom components&lt;/em&gt;. The base typography, spacing, color system, theme switching, and responsive behavior are already in place.&lt;/p&gt;

&lt;h3&gt;The Frontmatter Contract&lt;/h3&gt;

&lt;p&gt;Both formats share exactly the same frontmatter schema:&lt;/p&gt;

&lt;div class=&quot;code-comparison&quot;&gt;
  &lt;div&gt;
    &lt;span class=&quot;label&quot;&gt;Markdown article&lt;/span&gt;
    &lt;pre&gt;&lt;code&gt;---
title: My Article Title
date: 2026-02-22
summary: A short description.
status: published
tags: web, css, javascript
heroImageDark: /img/blog/slug/hero-dark-960w.jpg
heroImageLight: /img/blog/slug/hero-light-960w.jpg
cardImageDark: /img/blog/slug/hero-dark-og.jpg
cardImageLight: /img/blog/slug/hero-light-og.jpg
imageAlt: Descriptive alt text.
---

# My Article Title

Article body in **Markdown**...&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
  &lt;div&gt;
    &lt;span class=&quot;label&quot;&gt;HTML article&lt;/span&gt;
    &lt;pre&gt;&lt;code&gt;---
title: My Article Title
date: 2026-02-22
summary: A short description.
status: published
tags: web, css, javascript
heroImageDark: /img/blog/slug/hero-dark-960w.jpg
heroImageLight: /img/blog/slug/hero-light-960w.jpg
cardImageDark: /img/blog/slug/hero-dark-og.jpg
cardImageLight: /img/blog/slug/hero-light-og.jpg
imageAlt: Descriptive alt text.
---

&amp;lt;style&amp;gt;
  /* Custom overrides only */
&amp;lt;/style&amp;gt;

&amp;lt;h1&amp;gt;My Article Title&amp;lt;/h1&amp;gt;

&amp;lt;p&amp;gt;Article body in &amp;lt;strong&amp;gt;HTML&amp;lt;/strong&amp;gt;...&amp;lt;/p&amp;gt;&lt;/code&gt;&lt;/pre&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;p&gt;The content loader&apos;s flat frontmatter parser handles both identically: it splits on the &lt;code&gt;---&lt;/code&gt; delimiters, parses key-value pairs, and returns the same data structure regardless of what follows.&lt;/p&gt;

&lt;img src=&quot;/img/blog/blog-system-ux-and-media-pipeline/illustration-1-dark-960w.jpg&quot; alt=&quot;Split-view diagram showing Markdown and HTML content files flowing through frontmatter extraction into a unified article shell&quot; width=&quot;960&quot; height=&quot;640&quot; loading=&quot;lazy&quot; /&gt;

&lt;h2&gt;The Responsive Media Pipeline&lt;/h2&gt;

&lt;p&gt;Every article, whether Markdown or HTML, benefits from the same image delivery system. Source illustrations (generated using AI image tools and curated to match the site&apos;s prismatic color palette) are processed through a Python pipeline into optimized responsive variants.&lt;/p&gt;

&lt;h3&gt;From Source to Delivery&lt;/h3&gt;

&lt;p&gt;Each illustration starts as a high-resolution PNG with a transparent background in our site&apos;s prismatic color palette. A single command processes them:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;python3 scripts/blog/process-blog-images.py &amp;lt;article-slug&amp;gt;&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Each source image is a single transparent PNG. The pipeline composites it onto dark and light backgrounds, then encodes at 4 responsive widths in both WebP and JPEG, producing up to 12 responsive output files per source:&lt;/p&gt;

&lt;div style=&quot;margin: 1.5rem 0;&quot;&gt;
  &lt;div class=&quot;width-variants&quot;&gt;
    &lt;span class=&quot;width-pill&quot;&gt;1536w&lt;/span&gt;
    &lt;span class=&quot;width-pill&quot;&gt;1280w&lt;/span&gt;
    &lt;span class=&quot;width-pill&quot;&gt;960w&lt;/span&gt;
    &lt;span class=&quot;width-pill&quot;&gt;640w&lt;/span&gt;
  &lt;/div&gt;
  &lt;div style=&quot;display: flex; gap: 0.5rem; margin-top: 0.5rem;&quot;&gt;
    &lt;span class=&quot;format-badge webp&quot;&gt;WebP 82q&lt;/span&gt;
    &lt;span class=&quot;format-badge jpg&quot;&gt;JPEG 84q&lt;/span&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;p&gt;Hero images also get a social sharing crop at 1200×630 for Open Graph cards: one neutral WebP plus dark and light JPEG variants, adding 3 files. A single hero source produces &lt;strong&gt;15 optimized files&lt;/strong&gt;. Three inline illustrations add 36 more, bringing this article&apos;s total to &lt;strong&gt;51 responsive assets from 4 source images&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;Browser Delivery&lt;/h3&gt;

&lt;p&gt;The blog uses the &lt;code&gt;&amp;lt;picture&amp;gt;&lt;/code&gt; element with &lt;code&gt;&amp;lt;source&amp;gt;&lt;/code&gt; for WebP and &lt;code&gt;&amp;lt;img&amp;gt;&lt;/code&gt; fallback for JPEG. Both carry full &lt;code&gt;srcset&lt;/code&gt; and &lt;code&gt;sizes&lt;/code&gt; attributes computed from the actual CSS layout math:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;&amp;lt;picture&amp;gt;
  &amp;lt;source type=&quot;image/webp&quot;
          srcset=&quot;hero-1536w.webp 1536w,
                  hero-1280w.webp 1280w,
                  hero-960w.webp 960w,
                  hero-640w.webp 640w&quot;
          sizes=&quot;(min-width: 2560px) 1536px,
                 (min-width: 1920px) 1280px,
                 (min-width: 1440px) 979px,
                 (min-width: 1088px) 90vw, 100vw&quot;&amp;gt;
  &amp;lt;img src=&quot;hero-dark-960w.jpg&quot;
       srcset=&quot;hero-dark-1536w.jpg 1536w, ...&quot;
       width=&quot;960&quot; height=&quot;640&quot;
       alt=&quot;...&quot;&amp;gt;
&amp;lt;/picture&amp;gt;&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;WebP sources are theme-neutral: the same file serves both dark and light modes because WebP preserves the original transparent-background illustration. JPEG sources carry the theme suffix (&lt;code&gt;hero-dark-960w.jpg&lt;/code&gt;) because JPEG requires an opaque background. The browser picks the optimal width variant automatically. On a 1440px viewport, it loads the 960w image. On a 4K display, it loads the 1536w. Mobile devices get the 640w. No wasted bandwidth, no blurry scaling.&lt;/p&gt;

&lt;h3&gt;Theme-Aware Image Switching&lt;/h3&gt;

&lt;p&gt;Both hero and inline images participate in theme switching. When the user toggles between dark and light mode, JavaScript swaps the &lt;code&gt;src&lt;/code&gt; attribute on every themed image. Inline article images use a naming convention (&lt;code&gt;illustration-1-dark-960w.jpg&lt;/code&gt; → &lt;code&gt;illustration-1-light-960w.jpg&lt;/code&gt;) that the loader detects automatically and applies &lt;code&gt;data-theme-image-dark&lt;/code&gt; / &lt;code&gt;data-theme-image-light&lt;/code&gt; attributes for switching.&lt;/p&gt;

&lt;img src=&quot;/img/blog/blog-system-ux-and-media-pipeline/illustration-2-dark-960w.jpg&quot; alt=&quot;Responsive image pipeline visualization: a single source PNG branching into WebP and JPEG variants at four width breakpoints&quot; width=&quot;960&quot; height=&quot;640&quot; loading=&quot;lazy&quot; /&gt;

&lt;h2&gt;Why Both Formats?&lt;/h2&gt;

&lt;p&gt;Markdown is right for most content. It&apos;s fast to write, hard to break, and the parser handles security (escaping HTML entities before rendering). For a dev blog where the content is primarily prose with code snippets, Markdown removes friction.&lt;/p&gt;

&lt;p&gt;But some articles are about the web itself: demonstrating CSS techniques, showing interactive components, or (like this one) using custom layouts that make the content more effective. For those, writing HTML with inline styles is more natural and more capable than fighting Markdown&apos;s limitations.&lt;/p&gt;

&lt;p&gt;The dual-source approach means we never have to choose. Each article uses whichever format serves its content best, and the reader never knows the difference.&lt;/p&gt;

&lt;h2&gt;Implementation Details&lt;/h2&gt;

&lt;p&gt;For developers interested in the mechanics, here&apos;s how the content pipeline branches:&lt;/p&gt;

&lt;table class=&quot;architecture-table&quot;&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Stage&lt;/th&gt;
      &lt;th&gt;Markdown (.md)&lt;/th&gt;
      &lt;th&gt;HTML (.html)&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;File registration&lt;/td&gt;
      &lt;td colspan=&quot;2&quot; style=&quot;text-align: center;&quot;&gt;index.json (same for both)&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Frontmatter extraction&lt;/td&gt;
      &lt;td colspan=&quot;2&quot; style=&quot;text-align: center;&quot;&gt;splitFrontmatter() (same parser)&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Body processing&lt;/td&gt;
      &lt;td&gt;MarkdownParser.parse()&lt;/td&gt;
      &lt;td&gt;Direct injection (no parsing)&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Security model&lt;/td&gt;
      &lt;td&gt;All content escaped by parser&lt;/td&gt;
      &lt;td&gt;Author-trusted (same as any page)&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Inline styles&lt;/td&gt;
      &lt;td&gt;Not supported&lt;/td&gt;
      &lt;td&gt;&amp;lt;style&amp;gt; tags preserved&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Inline scripts&lt;/td&gt;
      &lt;td&gt;Escaped by parser&lt;/td&gt;
      &lt;td&gt;&amp;lt;script&amp;gt; tags preserved but not executed&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Theme switching&lt;/td&gt;
      &lt;td colspan=&quot;2&quot; style=&quot;text-align: center;&quot;&gt;Automatic (same JS handles both)&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SEO / social cards&lt;/td&gt;
      &lt;td colspan=&quot;2&quot; style=&quot;text-align: center;&quot;&gt;Identical (driven by frontmatter)&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The security model difference is worth noting. Markdown articles are sanitized: the parser escapes HTML, preventing injection even if the source file were compromised. HTML articles are trusted, like any other page on the site. This is the right tradeoff: HTML articles are authored by the site owner, not user-generated content.&lt;/p&gt;

&lt;img src=&quot;/img/blog/blog-system-ux-and-media-pipeline/illustration-3-dark-960w.jpg&quot; alt=&quot;Implementation comparison table showing how Markdown and HTML articles flow through different processing paths in the content pipeline&quot; width=&quot;960&quot; height=&quot;640&quot; loading=&quot;lazy&quot; /&gt;

&lt;h2&gt;What&apos;s Next&lt;/h2&gt;

&lt;p&gt;The foundation is in place: two content formats, one pipeline, responsive media at every breakpoint. Future improvements might include:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;AVIF support&lt;/strong&gt;: adding a third format tier for browsers that support it, further reducing file sizes&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Interactive HTML articles&lt;/strong&gt;: embedding live code playgrounds, interactive diagrams, or data visualizations&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Automated content validation&lt;/strong&gt;: CI checks that verify frontmatter completeness and image pipeline output for every article&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal isn&apos;t complexity for its own sake. It&apos;s having the right tool available when the content demands it, while keeping the default path simple and fast.&lt;/p&gt;
</content></entry><entry><title>Agentic IDE Workflows for Enterprise Teams</title><id>https://benlive.tv/blog/agentic-ide-workflows-for-enterprise-teams/</id><link href="https://benlive.tv/blog/agentic-ide-workflows-for-enterprise-teams/"/><updated>2026-02-14T12:00:00-05:00</updated><content type="html">&lt;h1&gt;Agentic IDE Workflows for Enterprise Teams&lt;/h1&gt;
&lt;p&gt;Most teams don&apos;t fail at AI adoption on install day. They fail around week three, when four engineers using the same assistant start producing four different kinds of output, and nobody can say which one matches the team&apos;s standard.&lt;/p&gt;
&lt;p&gt;This is the operating model I use for IDE-first work at my day job, where I&apos;m a QA automation engineer on a government software contract and the person who ended up owning how the team uses Amazon Q Developer. Everything here is about reliability over novelty, because on a regression suite that gates releases, novelty is not the goal.&lt;/p&gt;
&lt;h2&gt;Install is step zero&lt;/h2&gt;
&lt;p&gt;After the extension is installed and signed in against the organization&apos;s identity, I run a short acceptance check before anyone does real work:&lt;/p&gt;
&lt;ol&gt;&lt;li&gt;Confirm the identity context is the enterprise one, not a personal account&lt;/li&gt;&lt;li&gt;Confirm which workspaces the assistant can see and what the data handling expectations are&lt;/li&gt;&lt;li&gt;Run one non-sensitive smoke prompt&lt;/li&gt;&lt;li&gt;Record the extension version in the onboarding doc&lt;/li&gt;&lt;/ol&gt;
&lt;p&gt;The smoke prompt is boring on purpose:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;text&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;Summarize this file in five bullets and propose one deterministic test improvement.&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;If that loop isn&apos;t stable, nothing else gets attempted. I&apos;ve seen more time lost to a misconfigured sign-in than to any model limitation.&lt;/p&gt;
&lt;h2&gt;Ask mode, then agent mode, then ask mode again&lt;/h2&gt;
&lt;p&gt;The mode toggle in Q Developer is scope control, and I use it as a sequence rather than a preference.&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;Ask&lt;/strong&gt; to understand the current state and the options. No edits yet.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Agent&lt;/strong&gt; to apply the scoped change with explicit constraints.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Ask&lt;/strong&gt; again to review the diff for risk and to list what should be verified.&lt;/li&gt;&lt;li&gt;Run the tests and attach the outcome to the pull request.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;Ask mode keeps me from committing to a change I haven&apos;t thought through. Agent mode keeps me from spending the afternoon typing boilerplate. Going back to ask mode at the end is the step people skip, and it&apos;s the one that catches the edit that touched a file it shouldn&apos;t have.&lt;/p&gt;
&lt;h2&gt;Context packs, version controlled&lt;/h2&gt;
&lt;p&gt;Prompt quality is a system problem, not a per-message problem. The single biggest reduction in output variance across the team came from writing the context down and checking it in.&lt;/p&gt;
&lt;p&gt;I keep two kinds of files. The personal ones are portable behavior constraints, the way I want any assistant to work regardless of repo:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;md&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-md&quot;&gt;Role: senior engineer optimizing for production safety.
Priorities:
- Regression risk before style feedback.
- Scope stays inside the stated task.
- Every response ends with a verification summary.&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The repository ones are engineering rules the whole team owns, reviewed in pull requests like any other change:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;md&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-md&quot;&gt;- Behavior changes require the matching automated test update.
- No modifications to files outside the stated task.
- Prefer deterministic assertions over timing-based waits.
- Non-trivial releases include rollback notes.&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;At work these live as the Rules, Profiles, and Prompts the team uses daily. On my own repos the same idea shows up as &lt;code&gt;CLAUDE.md&lt;/code&gt; and &lt;code&gt;AGENTS.md&lt;/code&gt;. The format matters less than the discipline: the file is the standard, the standard has a diff history, and the diff history has reviewers.&lt;/p&gt;
&lt;h2&gt;Model routing as an explicit decision&lt;/h2&gt;
&lt;p&gt;The IDE gives you a model picker, and the default habit is to set it once and forget it. I treat it as a per-task decision with a written policy: start on the fastest model likely to pass the quality gates, escalate only when a real failure shows up, and note when escalation happens. The escalation log tells you where the routing policy needs adjusting far better than any benchmark table does.&lt;/p&gt;
&lt;p&gt;I&apos;m deliberately not listing model names here. They change quarterly. The policy doesn&apos;t.&lt;/p&gt;
&lt;h2&gt;Quality gates that don&apos;t care who wrote the code&lt;/h2&gt;
&lt;p&gt;The same gates apply whether a person or an assistant produced the change:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;Scope gate.&lt;/strong&gt; The changed files match the declared task.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Test gate.&lt;/strong&gt; The relevant checks ran and the results are in the PR.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Documentation gate.&lt;/strong&gt; Behavior-affecting diffs update the docs that describe the behavior.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Risk gate.&lt;/strong&gt; Rollout and rollback are both clear.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;Making those explicit up front has saved more time than any model upgrade. It also removed a category of argument. Nobody has to decide whether AI-generated code deserves extra scrutiny, because all code gets the same scrutiny.&lt;/p&gt;
&lt;h2&gt;What changed on the team&lt;/h2&gt;
&lt;p&gt;The measurable part is the weekly regression cycle. Triage, diagnosis, and resolution got faster and more consistent once the shared prompt patterns existed, because a teammate diagnosing a flaky test starts from the same context I would. Script maintenance benefits the same way. The part that&apos;s harder to measure is that new people ramp faster, which is what pushed me to build an internal documentation app for onboarding, a story for another post.&lt;/p&gt;
&lt;p&gt;None of this required the newest model. It required treating the assistant&apos;s context like configuration, the assistant&apos;s output like any other contributor&apos;s, and the mode toggle like the scope control it is.&lt;/p&gt;</content></entry><entry><title>A CLI Agent Harness Playbook</title><id>https://benlive.tv/blog/cli-agent-harness-playbook/</id><link href="https://benlive.tv/blog/cli-agent-harness-playbook/"/><updated>2026-02-03T12:00:00-05:00</updated><content type="html">&lt;h1&gt;A CLI Agent Harness Playbook&lt;/h1&gt;
&lt;p&gt;When I need a change I&apos;d stake a deploy on, I stay in the terminal. Inspect, patch, test, summarize. It&apos;s the tightest feedback loop I&apos;ve found for AI-assisted work, and it&apos;s the one that produces diffs I&apos;m willing to merge without re-reading every line.&lt;/p&gt;
&lt;p&gt;This is how I run Claude Code and Codex day to day, and what I&apos;ve written into the repos so the discipline doesn&apos;t depend on my memory.&lt;/p&gt;
&lt;h2&gt;The contract for every run&lt;/h2&gt;
&lt;p&gt;Five steps. When I skip one, I usually pay for it later.&lt;/p&gt;
&lt;ol&gt;&lt;li&gt;&lt;strong&gt;Scope, with non-goals.&lt;/strong&gt; Writing down what I&apos;m not doing matters as much as what I am. &quot;Fix the rate limiter&apos;s memory growth. Do not touch the Firestore-backed path.&quot;&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Read the minimum.&lt;/strong&gt; Broad reads produce broad edits. If the task is one function, the assistant reads that function and its callers, not the directory.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;One behavior change per patch.&lt;/strong&gt; Bundled changes hide the one that broke something.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Targeted tests first, then wider.&lt;/strong&gt; The spec that covers the change, then the canary, then the full batch only when the change earns it.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;A handoff note.&lt;/strong&gt; Files changed, what was verified, what wasn&apos;t. Without this the run isn&apos;t finished, it&apos;s just stopped.&lt;/li&gt;&lt;/ol&gt;
&lt;p&gt;I turned the parts of that loop that kept getting skipped into Claude Code skills. &lt;code&gt;/diagnose&lt;/code&gt; forces a written, ranked hypothesis list before any fix. &lt;code&gt;/implement-plan&lt;/code&gt; picks up an existing plan document and executes it phase by phase instead of writing a new plan. &lt;code&gt;/audit&lt;/code&gt; runs a CSS, JavaScript, and HTML checklist over the diff before a commit. Those aren&apos;t clever prompts. They&apos;re the steps I noticed myself rationalizing away at the end of a long session.&lt;/p&gt;
&lt;h2&gt;Routing between Codex and Claude Code&lt;/h2&gt;
&lt;p&gt;I don&apos;t have a benchmark-driven rule for this. I have a task-shape rule.&lt;/p&gt;
&lt;div class=&quot;table-scroll-wrapper&quot;&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Task&lt;/th&gt;&lt;th&gt;Where it starts&lt;/th&gt;&lt;th&gt;When I switch&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Small fix plus its test&lt;/td&gt;&lt;td&gt;Codex&lt;/td&gt;&lt;td&gt;The root cause turns out to be architectural&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Multi-file refactor&lt;/td&gt;&lt;td&gt;Claude Code, plan first&lt;/td&gt;&lt;td&gt;Execution is the bottleneck, not understanding&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Release readiness review&lt;/td&gt;&lt;td&gt;Claude Code&lt;/td&gt;&lt;td&gt;Findings are tooling-heavy and need patching&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Script and tooling upkeep&lt;/td&gt;&lt;td&gt;Codex&lt;/td&gt;&lt;td&gt;The policy question is ambiguous&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Escalation is a response to a concrete failure, not a reflex. If the first choice is working, I don&apos;t upgrade for the sake of it.&lt;/p&gt;
&lt;h2&gt;One source of truth for two assistants&lt;/h2&gt;
&lt;p&gt;When more than one assistant touches the same repo, they have to read the same rules. The pattern I&apos;ve settled on is that the harness-specific files are adapters, not forks of the policy.&lt;/p&gt;
&lt;p&gt;On &lt;a href=&quot;https://github.com/benmcnulty/good-vibes&quot; target=&quot;&lt;em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;good-vibes&lt;/a&gt; that looks like @@INLINE&lt;/em&gt;CODE_0@@ for Codex, &lt;code&gt;CLAUDE.md&lt;/code&gt; for Claude Code, and &lt;code&gt;.github/copilot-instructions.md&lt;/code&gt; for Copilot, each stating that assistant&apos;s role and pointing at the shared roadmap in &lt;code&gt;docs/&lt;/code&gt;. On this site, &lt;code&gt;CLAUDE.md&lt;/code&gt; is the canonical file and &lt;code&gt;AGENTS.md&lt;/code&gt; is deliberately short: it tells any other agent to go read &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;I review changes to those files as carefully as I review code. Small drift between them is how you end up with one assistant running the full Playwright suite in a single process while the other one batches it, and then arguing with the results.&lt;/p&gt;
&lt;h2&gt;Validating paid APIs without burning credit&lt;/h2&gt;
&lt;p&gt;The chat on this site depends on OpenRouter, and before a functions deploy I want to know the model chain is alive. The gate is deliberately cheap:&lt;/p&gt;
&lt;div class=&quot;code-block&quot; data-code-block&gt;&lt;div class=&quot;code-block-meta&quot;&gt;&lt;span class=&quot;code-language&quot;&gt;json&lt;/span&gt;&lt;button class=&quot;code-copy-button&quot; type=&quot;button&quot; data-code-copy-button data-default-label=&quot;Copy&quot; aria-label=&quot;Copy code block&quot;&gt;Copy&lt;/button&gt;&lt;/div&gt;&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &quot;messages&quot;: [{ &quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;Reply with: ok&quot; }],
  &quot;promptKey&quot;: &quot;aiLab&quot;,
  &quot;model&quot;: &quot;&amp;lt;one model from the chain&amp;gt;&quot;
}&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;One probe per model. No long prompts in a health check. No blind retries. Every failure reason logged. The predeploy spec reads the model list from the same config file the function uses. A gate with its own hard-coded model IDs is a gate that goes stale the day a provider retires one, and free-tier catalogs retire models without asking.&lt;/p&gt;
&lt;h2&gt;Local models in the loop&lt;/h2&gt;
&lt;p&gt;Bounded tasks are where local models earn their place: offline scaffolding, low-risk text transforms, a first draft of documentation. &lt;a href=&quot;https://github.com/benmcnulty/localcrew&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Local Crew&lt;/a&gt; is the terminal-first version of that idea taken all the way. It pools Ollama and OpenAI-compatible endpoints across the machines on my network, routes tasks by resource tier, and runs the same inspect, patch, verify loop across devices instead of on one.&lt;/p&gt;
&lt;p&gt;I still don&apos;t let a local model be the final authority on a high-risk multi-file change. But &quot;final authority&quot; is a small fraction of the work.&lt;/p&gt;
&lt;h2&gt;Failure patterns that shaped the playbook&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Scope creep in autonomous runs.&lt;/strong&gt; The fix was a changed-file check before the test suite runs. If the diff is wider than the task, stop and ask.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Green tests, wrong behavior.&lt;/strong&gt; Assertions that were too loose. The fix was tying tests to the acceptance criteria in the plan, not to whatever the code happened to do.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Strong reasoning, weak execution.&lt;/strong&gt; Splitting planning and patching into separate runs with a checkpoint between them. Let each tool do the part it&apos;s good at.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Silent skips.&lt;/strong&gt; The one I watch for most. A runner that treats a killed process as a pass, or a suite that skips itself when a dependency is slow to start, looks green right up until it matters. The fix is a runner that&apos;s honest about exit codes and a readiness probe that retries before it gives up.&lt;/p&gt;
&lt;h2&gt;What holds up&lt;/h2&gt;
&lt;p&gt;The advantage of a CLI harness isn&apos;t the model inside it. It&apos;s that the loop is auditable. Every step leaves something a reviewer can read: a plan, a diff, a test result, a note. As models improve, the teams that enforce that contract will compound the gains. The teams chasing leaderboards will keep re-reading diffs.&lt;/p&gt;</content></entry><entry><title>Models and Harnesses</title><id>https://benlive.tv/blog/models-and-harnesses/</id><link href="https://benlive.tv/blog/models-and-harnesses/"/><updated>2026-01-24T12:00:00-05:00</updated><content type="html">&lt;h1&gt;Models and Harnesses&lt;/h1&gt;
&lt;p&gt;Most tool debates start with the model. I get better results starting with the harness.&lt;/p&gt;
&lt;p&gt;By harness I mean the thing wrapped around the model: how it reads the repo, how it applies edits, how it runs tests, what it hands back when it&apos;s done. A capable model inside a sloppy harness still produces changes I can&apos;t trust. I&apos;ve watched that happen enough times that harness choice is now the first decision, and model choice is the second.&lt;/p&gt;
&lt;h2&gt;What I evaluate a harness on&lt;/h2&gt;
&lt;p&gt;Four things, and if any one of them is weak the model doesn&apos;t rescue it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Context hygiene.&lt;/strong&gt; Can I keep what the assistant sees tight and repeatable? The harness needs to read the project&apos;s rules file on every session without me pasting it. On this site that file is &lt;code&gt;CLAUDE.md&lt;/code&gt;, and on my other repos it&apos;s usually &lt;code&gt;AGENTS.md&lt;/code&gt; next to it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Change control.&lt;/strong&gt; Are the edits easy to review and hard to sprawl? I want a diff that matches the task I described, not a diff that also &quot;cleaned up&quot; three unrelated files.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Verification loop.&lt;/strong&gt; Can the harness run the tests and report the result in the same session? If I have to leave the tool to find out whether the change works, I&apos;ll eventually skip that step, and so will anyone else.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Handoff quality.&lt;/strong&gt; When it&apos;s done, do I get a summary I could paste into a pull request? Changed files, what was verified, what wasn&apos;t.&lt;/p&gt;
&lt;h2&gt;The harnesses I actually use&lt;/h2&gt;
&lt;p&gt;I use four, and I use them for different things. This isn&apos;t a ranking.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt; is where I go when I need the codebase understood before anything gets touched. Planning a refactor, tracing a bug across a serverless function and the client that calls it, reading a suite of Playwright specs to see what they really assert. It&apos;s the harness behind most of the larger changes on this site, and it&apos;s the one that respects a plan file well enough that I can approve the plan and walk away.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Codex&lt;/strong&gt; does well in a scoped terminal loop: inspect a few files, patch, run the targeted test, summarize. When the task is &quot;make this one behavior change and update its test,&quot; Codex tends to stay inside the lines.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GitHub Copilot&lt;/strong&gt; is inline acceleration. Fixtures, repetitive edits, the third page object that looks like the first two. I don&apos;t plan with it. On &lt;a href=&quot;https://github.com/benmcnulty/good-vibes&quot; target=&quot;&lt;em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;good-vibes&lt;/a&gt; I gave it one job, reviewing commit diffs for style and consistency, and wrote that down in @@INLINE&lt;/em&gt;CODE_0@@ so the role stuck.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Amazon Q Developer&lt;/strong&gt; is what I use at work, in Visual Studio and VS Code, on a codebase that isn&apos;t mine to publish. The mode split between asking and acting is the part I lean on most, and I&apos;ll write more about that separately.&lt;/p&gt;
&lt;p&gt;That good-vibes setup is worth spelling out because it&apos;s the clearest small example of the idea: three assistants, three written roles. Codex owned planning and the roadmap in &lt;code&gt;AGENTS.md&lt;/code&gt;. Claude Code owned implementation and tests in &lt;code&gt;CLAUDE.md&lt;/code&gt;. Copilot owned review. None of them had to guess what the others were for.&lt;/p&gt;
&lt;h2&gt;Where the model still matters&lt;/h2&gt;
&lt;p&gt;Once the harness is settled, the model is a capability and cost decision.&lt;/p&gt;
&lt;p&gt;The chat on this site makes that concrete. Visitors see three tiers, Fast, Balanced, and Thoughtful. Behind those are three free-tier models on OpenRouter, and behind each of those is a local Ollama model I can run on my own machine with roughly the same character. The mapping lives in one place, &lt;code&gt;functions/model-config.js&lt;/code&gt;, and the rule is that every production model has a local stand-in. If OpenRouter retires a model, and it has, the chain falls through to the next one and eventually to a managed free router. I&apos;d rather answer a question with a smaller model than show an error.&lt;/p&gt;
&lt;p&gt;That&apos;s the whole model-selection philosophy in one feature: pick the cheapest model that clears the quality bar for the task, keep a fallback, and make sure I can reproduce the behavior locally before I ship it.&lt;/p&gt;
&lt;h2&gt;Where local models fit now&lt;/h2&gt;
&lt;p&gt;Local models are useful for bounded work. Prompt drafting, which is what &lt;a href=&quot;https://github.com/benmcnulty/promptpad&quot; target=&quot;&lt;em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Promptpad&lt;/a&gt; does against Ollama. Generating test data without sending anything to a third party, which is what &lt;a href=&quot;https://github.com/benmcnulty/synthetic-data&quot; target=&quot;&lt;/em&gt;blank&quot; rel=&quot;noopener noreferrer&quot;&gt;synthetic-data&lt;/a&gt; does from a browser tab. Enhancing a terse task description before it goes to a coding agent, which is what &lt;a href=&quot;https://github.com/benmcnulty/hopper&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;hopper&lt;/a&gt; does before it queues work for Claude Code or Codex.&lt;/p&gt;
&lt;p&gt;I still don&apos;t hand a local model a long multi-file change and accept the result on faith. The verification loop has to be tighter, and the scope has to be narrower. But &quot;interesting toy&quot; stopped being the right description a while ago. &lt;a href=&quot;https://github.com/benmcnulty/localcrew&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;Local Crew&lt;/a&gt; exists because I wanted to see how far a pool of ordinary machines on my own network could go, and the answer was further than I expected.&lt;/p&gt;
&lt;h2&gt;The metrics I keep&lt;/h2&gt;
&lt;p&gt;I stopped arguing about which model is best and started tracking things I can measure on my own work:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Time from prompt to a passing test&lt;/li&gt;&lt;li&gt;How many corrections I make before I&apos;d merge it&lt;/li&gt;&lt;li&gt;Regressions that show up after deploy&lt;/li&gt;&lt;li&gt;Whether the handoff note told a reviewer what they needed&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;A setup that looks impressive and fails those is a demo. A setup that passes them is a tool.&lt;/p&gt;
&lt;h2&gt;The short version&lt;/h2&gt;
&lt;p&gt;Treat the harness and the model as two decisions. The harness sets how reliable the workflow is. The model sets what it can do and what it costs. Tests and scope control decide whether any of it is mergeable. That order has held up for me better than chasing whichever model shipped last week.&lt;/p&gt;</content></entry><entry><title>Welcome to the Developer&apos;s Blog</title><id>https://benlive.tv/blog/intro-to-dev-blog/</id><link href="https://benlive.tv/blog/intro-to-dev-blog/"/><updated>2026-01-15T12:00:00-05:00</updated><content type="html">&lt;h1&gt;Welcome to the Developer&apos;s Blog&lt;/h1&gt;
&lt;p&gt;This is where I share real experiences from building AI-enhanced development tools. No hype, no magic. Just what works, what doesn&apos;t, and what it actually takes to ship reliable systems when LLMs are in the loop.&lt;/p&gt;
&lt;h2&gt;What to Expect&lt;/h2&gt;
&lt;p&gt;I&apos;ve spent the last decade building production software and the last few years pushing that work into AI-assisted workflows. The AI Lab is my playground for experiments, but I still care about the fundamentals: observability, resiliency, accessibility, and developer experience.&lt;/p&gt;
&lt;p&gt;You won&apos;t find:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Breathless predictions about AGI timelines&lt;/li&gt;&lt;li&gt;Tutorials that ignore production realities&lt;/li&gt;&lt;li&gt;Marketing speak disguised as technical content&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;You will find:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;Real problems and how I solved them (or didn&apos;t)&lt;/li&gt;&lt;li&gt;Technical deep dives when warranted&lt;/li&gt;&lt;li&gt;Honest assessments of tools and techniques&lt;/li&gt;&lt;/ul&gt;
&lt;h2&gt;Why Now?&lt;/h2&gt;
&lt;p&gt;The AI development space is moving fast, and it&apos;s easy to get lost in the noise. I want a place to capture what I&apos;m learning in a format that helps other engineers make grounded decisions.&lt;/p&gt;
&lt;p&gt;If you&apos;re building tools for humans (not just demos for a slide deck), we probably care about the same things.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/intro-to-dev-blog/illustration-1-dark-960w.jpg&quot; alt=&quot;A developer&apos;s desk with AI assistant interfaces, code review panels, and testing dashboards arranged in a productive workflow&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;h2&gt;Topics I&apos;ll Cover&lt;/h2&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;AI-enhanced development workflows&lt;/strong&gt;: What actually works in production&lt;/li&gt;&lt;li&gt;&lt;strong&gt;System architecture&lt;/strong&gt;: Designing for reliability when LLMs are in the loop&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Testing strategies&lt;/strong&gt;: How to validate AI-assisted code&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Developer experience&lt;/strong&gt;: Building tools people actually want to use&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Team practice&lt;/strong&gt;: How to ship responsibly with new tools, not just quickly&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;&lt;img src=&quot;/img/blog/intro-to-dev-blog/illustration-2-dark-960w.jpg&quot; alt=&quot;Connected network of blog topics: architecture diagrams, test suites, developer tools, and team workflows linked by glowing data streams&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;/p&gt;
&lt;h2&gt;Let&apos;s Connect&lt;/h2&gt;
&lt;p&gt;If you&apos;re working on similar problems or have questions about anything I write here, reach out. This blog is meant to be a conversation, not a broadcast. I&apos;ll share what I&apos;m building, and I&apos;ll be just as interested in what&apos;s working for you.&lt;/p&gt;
&lt;p&gt;Looking forward to sharing what I&apos;m learning as I build.&lt;/p&gt;
&lt;p&gt;Ben&lt;/p&gt;</content></entry></feed>