Originally published as an Advisory Hour Substack post on 2026-08-15. This first-party site copy preserves the text, localizes public supporting media, and connects the record to its project trail.
Original post

The gift of local video is that it allows low-cost experimentation. The experiments without local AI require a credit card. Up until this week, my video generation experiments used a third party API. The fire department showed up at my door asking what was smoking. I handed over my wallet, “My credit card?”
It doesn’t take long to reach the inflection point where API based video generation switches from “cool, neat trick” to “maybe I can skip eating out this week”. It’s the costs. On the higher-end of video generation over an API the costs are shockingly ridiculous. They’re also the kind of costs every SaaS player with an API wishes they could charge per API request.

You read that right. $1.25 per second of output video. Here’s the catch. You do a 5 second test run of an idea, blow $6.00 USD and at the same time have a video the video generation model didn’t quite handle all that great. Maybe an arm disappeared. Maybe instead of a medieval setting you got a 1973 romcom thriller. Either way? You’re out $6.00 faster than it goes at the bar during happy hour. Local AI changes the price to $0.00 + your electricity bill. That’s a great price! Now when a transforming robot video completely fails it’s totally fine. A whole lot better to spend $0.00 on something like this over $6.00.
I think it’s a robot having a bad encounter with a taco truck. Not that I’m exploring transforming robots and taco trucks. But if I were? It’s effectively free.
Downloading a Universe
It shouldn’t be possible to download an entire universe. It takes a couple hours to do so. It’s not even all that big. It takes less time to download an entire approximated universe than it does to get all my laundry caught up after a busy week. I can’t approximate clean clothes. Well, maybe my gym clothes-much to the chagrin of my wife.
See, I have this endless fascination with how video models compressed the known world. Seemingly these digital bits of wonder work with any time and place. With a mere prompt anyone on a keyboard may pull out of a video model all manner of understanding of the real world, even place.
To be clear, by place I mean a location the video model has approximated understanding of what the city you live in looks like. For example, I had the video model do a drive into the City of Des Moines. I wasn’t sure what to expect, but it looked a lot like driving I-80. The music is so generic and unexpected. Why does a video model so easily handle voice, music, and place? I get the technology aspects, but I don’t get the philosophical aspects. I downloaded a universe? Why is this possible?
And experimentally you can drive combinations of things. Having a video model that I can run to explore any approximated location is great. However, H3 supports image references which means I can pull a real world location and have it show me an imagined scene. Let’s do a bicycle ride in downtown Des Moines next.
This next video is based on a screen clip from google maps. The video model put music on as an overlay as go for a bike ride towards the downtown Principal building. The Principal building (or 801 Grand) of Des Moines is the tallest building in Iowa. That’s not saying much, as most of our buildings are corn silos. I often stare at this building in disbelief. Elon and the SpaceX team intend to ship one of these on the regular into outer space. We don’t have a robust space industry in Des Moines, but we do have bike lanes. Let’s go on that fictional ride in an imagined version of an approximated understanding of Des Moines.
Exploration of location also goes in weird directions-straight down. The next video takes us on a specific ride of “what’s it like to jump off a tall building” and uses fast draft settings available in a video model. There’s a step sequence you can set which helps guide for how polished the overall video is, at the trade for time. I wasn’t committing to the idea of jumping off the building as a fully done video. I was, however, morbidly curious about it.
It’s so weird. What surprised me was just how wonderfully weird this clip is and also the voices that can’t quite be understood. It reminded me of a dream sequence you might have-the classic “I’m falling” dream. I’ve watched the clip enough times to realize that I don’t understand at all what’s going on in it.
From videos in real world as fiction to just videos in fictional worlds…
This week I’ve found myself writing more than expected about Ashwicke. The Ashwicke began as a fun room-puzzle type game that has evolved into a collection of short video. It’s a strange place. Ashwicke is neither horror nor a tragedy experience. The Ashwicke is a mad mix of Victorian-era-esque scifi. There’s likely a better word for that as a concept. I just refer to it as the Ashwicke. I’ve grown fond of the place in a bizarre way, and if you’ve watched some of the other videos then you’ll appreciate the deep lore cut of this video, too.
The Sky Sticker

I asked myself if I could take the video of driving along interstate I-80 (fictional I-80) and get a fun sticker. What if every interstate mile marker has a sticker? Stickers are an interesting way to share a journey. You collect them at a time and place and they serve as a kind of adhesive ledger of that time in your life. This sticker is Bubble-free stickers - When the sky says "We need to talk" along those lines. Instead of a sticker at a particular mile marker, it’s the sky saying… “We need to talk”.
Connected work
- Local AI: Minimax H3 on a Nvidia 3080: Read the local video-generation experiment that pushed a liquid-cooled 3080 through heat and physics tests.
- Last Train to Sunset: Snapdragon Video Test: Follow an earlier hardware test that finally produced a useful local video breakthrough.
- Model Training: Browse local inference experiments, released models, and hardware notes.