Your AI Second Brain Is Slowly Rotting (Here's How to Fix It)
Demonstrates maintaining an AI second brain using the second-brain-audit skill from coleam00/skills to ensure stored facts remain true, distinguishing between state and events.
Videos
Find videos about LLM models, hardware, runtimes, benchmarks and practical AI workflows.
Demonstrates maintaining an AI second brain using the second-brain-audit skill from coleam00/skills to ensure stored facts remain true, distinguishing between state and events.
Demonstrates the Gauntlet Loop multi-agent pattern where a Lead Agent delegates to Designer, Tester, and UI sub-agents while a Judge Agent evaluates results against an external benchmark until passing
Tests 1-bit Bonsai 27B performance, memory, agency, OpenAI Human Eval, Kanban, Sand Physics, and Dungeon Crawler against ~90% FP16 intelligence retention claim.
Explains graph engineering using LangGraph, Google ADK and AutoGen, positioning it against prompt, context, harness and loop techniques while distinguishing nodes and addressing applicability costs.
Installs and tests Qwen3.6-27B-Fable-Fusion-711-NM-DAU-NEO-MAX-MTP-GGUF locally, an open source multi-stage model fine tune for consumer hardware.
Tests Ling-3.0-tiny for token efficiency and production-scale agentic inference capabilities.
Running DeepSeek V4 Flash locally on two NVIDIA DGX Spark units while developing BridgeMind One and updating BridgeBench using Claude Opus 5, Fable 5, GPT 5.6 SOL, Kimi K3, Claude Code, and Codex.
Tests Meta Muse Code and Muse Spark 1.2 on browser workflows, C++ games, multimodal CAD, frontend dev, 3D modeling, flight sims, cinematic games, and FPS development.
Locally installs and tests Audio8 TTS Preview, a 0.6B multilingual text-to-speech model featuring zero-shot voice cloning capabilities.
Tests Muse Spark 1.2 and Muse Code via World of AI Bench, real-world coding, WebGL, and agentic software engineering, comparing performance against GPT-5.6 Luna, GPT-5.5, Claude Opus 4.8, and Grok 4.5
Evaluates Boris Cherny's advice to delete CLAUDE.md, skills, and hooks using benchmarks on Archon, comparing results against Anthropic's reduction of Claude Code's system prompt for Opus 5.
Evaluates Inkling-Small MLX Q9 using Inferencer App on M3 Ultra 512 GiB hardware, highlighting its status as a benchmark-topping open weight model with multimodal inference capabilities.
Installs and tests Muse Code by Meta, a terminal coding agent designed to handle complete software engineering tasks.
Evaluates Grug 27B QAT Q4 against Qwen 27B Q4 via Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot tests on a 16GB system.
Explains loop engineering using ReAct agents, contrasting it with prompt, context, and harness engineering, covering second-model evaluators, self-prompting, and shared memory outside agent state.
Compares KAT-Coder-V2.5-Dev and Qwen3.6-35B-A3B within Pi Coding Agent on a .NET repo for codebase analysis and prioritized feature roadmaps.
Installs and tests Shieldstral from Mistral, a compact 3B-parameter, policy-adaptive multimodal safety classifier.
Regina test set results compare Gemma 26B and Qwen 35B using IQ4 quants from unsloth, identical sampling parameters, prompts, harness, and llama-server versions on Nvidia GPUs.
Installation guide for running MiniMax H3 in ComfyUI, covering text-to-video, image-to-video, reference-to-video, video editing, audio reference, performance optimization, low VRAM usage, and LoRA tra
Demonstrates connecting Higgsfield MCP server to Claude Code for an end-to-end cinematic workflow using MiniMax-H3, Seedance 2.0, Kling 3.0, Grok Imagine 1.5, Google/Gemini Omni Flash, Veo 3.1 and Ali
Locally installs and tests Nemotron-3-Embed-1B, a text embedding model optimized for retrieval and semantic similarity tasks.
Qwen 3.6 35B A3B and Gemma 4 26B A4B are tested on updating documentation and verifying tests for mature legacy code using the swing-extras project.
Tests Traycer multi-agent collaboration using Claude Fable 5 and local Qwen 27B on a coding task, comparing results against Qwen 27B solo performance.
Tests Qwen 3.8 Max in coding, frontend, Three.js, agents, research, and visual reasoning versus Claude Opus 5, GPT-5.6 Sol, and Gemini. Covers Qwen 3.8 27B open-weight release.
Compares three open stalwart models: Qwen 3.8 Max, DeepSeek V4 Flash, and Kimi K3 through tests evaluating their performance capabilities.
Tests Qwen 3.8-Max vision capabilities.
Speed test of DeepSeek-V4-Flash-0731-MLX using Inferencer App v2.2.3 and DSpark on Mac Studio M3 Ultra 512 GiB.
Sets up a local TTS server using dots.tts and Qwen3-TTS, then integrates both models into the TalkWithMe AI group chat application to compare their performance.
Tests Qwen3.8 Max across browser workflows, C++ game creation, 3D CAD modeling, cinematic games, frontend design, creative writing, and multimodal coding.
Evaluates KAT Coder V2.5 Dev against Qwen 3.6 35B A3B using performance, memory, agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot tests.
Tests Qwen 3.8-Max, described as the most capable model in the Qwen family to date.
Compares Seedance 2.5 and Minimax H3 across multimodal features, instruction following, sketch to animation, storyboard to commercial, UI motion graphics, music videos, language tests, specs and costs
Evaluates Elements Labs' Bionic for local, private coding and software engineering using open-weights models.
Evaluates local AI agent harnesses using an RTX 3060 and qwen36-35b-a3b-mtp, comparing OpenClaw and Hermes against a custom build with llama-server, Telegram integration, and n8n automations.
Locally installs and tests ARK-ASR-3B, a multilingual automatic speech recognition model hosted on Hugging Face under Audio8/ARK-ASR-3B.
Tests GPT-5.6 Luna performance on browser OS, C++ games, FPS dev, multimodal coding, 3D CAD models, watch frontends, Chrono City C++, HQ Street Yeet, and result improvement.
Tests HappyHorse 1.0 from Alibaba for generating cinematic frames.
Overview of Deepseek V4 Flash 0731, Seedance 2.5, Minimax H3, Kimi K3, Instella, Inkling Small, Prism, Wonder, Phi Zero, ID V2V, Crisper Whisper, Redesign, Ideogram object remover, Higgsfield, Gemini
Evaluates DeepSeek-V4-Flash-0731-MLX using Inferencer App v2.2.3 on M3 Ultra 512 GiB across local and cloud setups against GLM 5.2 benchmarks.
Compares MiniMax H3 against Google Gemini (Veo 3.1) and OpenAI's Sora 2 on cost, resolution, and generation quality using text, image, video, and audio inputs.
Tests DeepSeek-V4-Flash-0731 checkpoint for speed, coding quality, tool use, hallucinations and reliability, then compares results against Tencent HY3 and Qwen3.6-27B.
Installs and demonstrates using Obsidian with Ollama and hermes-agent for free hands-free note management.
Demonstrates a single-user DGX Station setup with 748GB of unified memory running GLM 5.2 in VSCode.
Voice features on ChatGPT and Claude enable hands-free building, editing, and iteration of interactive 3D designs, such as generating an Iron Man suit, scaling helmets, changing colors, and adding sol
Benchmarks DeepSeek V4 Flash against GPT-5.6 Luna, Kimi K3, GLM 5.2, and DeepSeek V4 Pro via WoAI Benchmark in coding, AI agents, frontend development, Three.js, game generation, and UI design.
Installs and tests Conversation Stenography locally using nethical6/conversation-steganography repository for private AI chats.
Covers Claude Opus 5 tests, Google Earth Nano Banana, Meta AI Agents, BUZZ, Grok Build Mode, Gemini features, Midjourney 8.2, Mirage Avatar X, HeyGen, LinkedIn AI button, Friend wearable, robot bans a
Tests DeepSeek V4 Flash on browser workflows, C++ games, FPS dev, frontend design, 3D CAD modeling, and the Steve’s PC Repair game test.
Demonstrates QVAC, an open-source local AI platform providing Whisper, embeddings, Qwen, and text-to-speech via a single npm install across multiple operating systems under Apache 2.0 license.
Evaluates Laguna S 2.1 118B A8B against its XS version using Performance, Memory, Agency, OpenAI Human Eval, Sand Physics, Kanban, Dungeon Crawler, Blender and Godot tests.