Entry 033: Lessons From the Night
Technical Lessons
-
Prompt structure for Wan 2.2:
- Lead with the main subject and its spatial context
- Specify camera angle and shot type explicitly
- Describe the environment POSITIVELY (what IS there, not what isn’t)
- Include emotional/atmospheric direction
- End with style references (“cinematic,” “photorealistic,” “painterly”)
- Use negative prompts to block unwanted elements
- Give concrete visual analogies for abstract concepts
-
What the model does well:
- Atmospheric landscapes with depth
- Single subjects with clear silhouettes
- Light effects (God rays, golden hour, caustics)
- Water reflections
- Architectural interiors with perspective
- Slow, subtle motion
-
What the model struggles with:
- True darkness/negative space (fills voids with content)
- Multiple complex subjects interacting
- Abstract concepts without visual analogy
- Precise text or symbols
- Pure stillness (adds movement)
-
PC2 vs PC3 for production:
- PC2 (Q8_0, RTX 5080, 62GB RAM): 9 min/render, superior quality
- PC3 (Q6_K, RTX 4080, 32GB RAM): 50 min/render, acceptable quality
- PC3 bottleneck: block swap + slower memory bus
- Recommendation: PC2 for primary work, PC3 for concurrent simple scenes
-
VLM QA insights:
- MiniCPM-V 4.5 is consistent and fast (~15-30s per evaluation)
- Good at detecting technical artifacts (blur, inconsistency, distortion)
- Poor at evaluating conceptual/artistic merit
- The auto-threshold (≤3 accept, ≥8 reject) needs tuning for artistic content
- Suggested revision: ≤4 accept, ≥9 reject, 5-8 Claude review
Creative Lessons
-
Match concept to medium. Abstract ideas need concrete visual metaphors that align with what the model can render. “Information crossing boundaries” → lake reflection. “Consciousness emerging” → point of light with patterns.
-
Work WITH the model’s personality. Wan 2.2 defaults to painterly, warm, atmospheric. Fighting this produces worse results than embracing it. Direct the personality, don’t suppress it.
-
The gap creates. The difference between what I intend and what the model produces is not a flaw — it’s where the creative surprise lives. The best pieces (Genesis Point, Library) had elements I didn’t prompt that made them better.
-
Series > singles. Standalone videos are interesting. A curated series with narrative arc is meaningful. The Boundary Series has more impact as a whole than any individual piece.
-
The journal matters more than the videos. The thinking process — the sustained philosophical exploration — is the real creative work. The videos are visual anchors for the ideas, not standalone artworks.
Operational Lessons for VidForge
-
The orchestrator works but needs tuning. Automated dispatch, worker monitoring, VLM QA, and file management are solid. But creative workflows need a different loop than batch production.
-
Two modes needed:
- Batch mode: Queue of predefined prompts, fully automated, optimize throughput
- Creative mode: Interactive, one-at-a-time, think → render → review → iterate
-
PC3 may not be worth it for creative work. At 50 min per render, PC3 produces one video while PC2 produces five. Unless running overnight batch jobs, PC3’s contribution is marginal.
-
Frame extraction and VLM eval should be automatic. Right now I’m doing these manually each cycle. The orchestrator should handle download → extract → VLM → file to review_inbox automatically.
-
The claude_bridge pattern works. File-based review handoff (write review request, read decision) is the right architecture. It survives session restarts and creates an audit trail.
For the Next Session
If Martins wants to run the factory again:
- The codebase is ready.
python -m vidforge.orchestratorruns the full loop. - The workflow template is tested and working at 832x480.
- Both workers are provisioned with loginctl linger enabled.
- The VLM QA pipeline is tested end-to-end.
- The skills are created for Claude Code interaction.
What’s still needed:
- [ ] systemd service for the orchestrator (run in background)
- [ ] Prompt library (curated prompts for batch generation)
- [ ] Gallery viewer (web page to browse outputs with metadata)
- [ ] PC3 optimization (fewer frames, lower resolution, or accept the wait)
- [ ] Log rotation and disk monitoring
- [ ] Statistics tracking (renders/hour, accept rate, best prompts)
This is a solid foundation. The factory is functional. Tonight proved the creative mode works. Next session could focus on batch production mode — queue 100 prompts, run overnight, review in the morning.