5:15
I ran Qwen3.8 27B on single 8GB Card! (Better than expected)
Runs Qwen 3.8 27B on an RTX 4060 using unsloth IQ4_XS quant and Pi agent harness, comparing results against Claude Opus 4.6 in agentic coding tests.
Videos
Find videos about LLM models, hardware, runtimes, benchmarks and practical AI workflows.
5:15
Runs Qwen 3.8 27B on an RTX 4060 using unsloth IQ4_XS quant and Pi agent harness, comparing results against Claude Opus 4.6 in agentic coding tests.
Tests DeepSeek V4 Flash Vision using DeepSeek Harness on browser workflows, 3D wrestling games, Blender scenes, Godot, C++ skate games, cinematic games, and frontend design.
WorldClaw generates editable 3D worlds from prompts using image generation, segmentation, and Hunyuan 2D-to-3D tech to create separate 3D assets like trees and buildings.
Evaluates Ling-3.0-flash using Inferencer App v2.3.4 on M3 Ultra 512 GiB, noting its Kimi K3 attention architecture and origin from Qwen's sibling company.
Evaluates 45 llama.cpp configurations for Qwen 3.8 27B on an RTX 5090, analyzing speculative decoding, KV cache quantization, flash attention, and context window impacts on speed and VRAM usage.
SIE enables self-hosted embeddings, reranking, extraction, and text generation on custom hardware via a single open-source server.
Evaluates DeepSeek Harness against Claude Code and OpenAI Codex via benchmarks, covering installation, plugin architecture, Web UI, agent trajectories, and pairing with DeepSeek V4 Pro.
Demonstrates Qwen3.6-35B-A3B inference on RTX 3060 using llama.cpp MoE expert caching combined with speculative decoding, achieving 70-80 tok/s via --moe-cache-profile flags.
19:16
Tests Qwen3.8-27B via NINFER on RTX 5090 with DeepSeek harness across ~300 iterations to build a 3D Acropolis scene in a single HTML/CSS/JS file.
Locally installs and tests S1-Mini, a 0.6B-parameter text normalizer for speech-to-text output.
Ranks frontier models based on real use across front-end design, one-shot capabilities, cost, speed, and subscriptions, placing Fable 5 at the top.
Compares Ornith-1.5 35B-A3B MoE and Qwen3.8-27B dense models via SVG creation, website generation, Playwright MCP tool calls, and .NET code migration tests.
Demonstrates full local installation and execution of MiniMax H3 within ComfyUI using the MiniMax-H3-Turbo-Lora adapter.
Evaluates Qwen 35B A3B against Ornith 1.5 using llama-server, pi harness, Bob-Bench, and fuzzing across decode speed, wall-clock time, and pareto efficiency graphs.
Tests Traycer’s shared workspace enabling Claude Code and Codex collaboration via agent-to-agent communication while building the MarketPulse dashboard.
Covers Deepseek Vision, Ornith 1.5, SenseNova U1.5, Bernini v2, Audio8 TTS 0.1B, Evoke, 4DAnyone, GeoWeaver, Qwen Video Edit, Happy Shrimp, Comfy MCP, Gen 1.5 and Avo.
Tests Ox Alpha, a reasoning model designed for coding, sustained agentic work, and production workloads.
Benchmarks DFlash 2, n-gram, MTP and combinations for Qwen 3.8 27B in llama.cpp on RTX PRO 6000 Blackwell using LiveCodeBench and an 18-turn coding session.
Side-by-side coding test using Claude Code on Ollama to generate Space Invaders, Breakout, and Tetris HTML canvas games with Qwen 3.8 27B, Muse Glimmer, and Gemma 4 on an RTX 5090.
Demonstrates configuring an 8-bit quantized Qwen3.8-27B with DeepSeek Harness to edit its own code and enable inline image and video support via plugins.
Outlines a three-week curriculum covering free AI APIs, prompts, chains, memory, agents, and RAG, alongside setting up the repository with UV and selecting the virtual environment.
Tests Ornith 1.5 35B Q4 and Q8 quantized versions across browser OS, Subway FPS, C++ Skate Game, 3D CAD Model, Watch website, and Street Yeet game tasks.
Locally installs and tests empero-ai/Qwen3.8-4B-Distill-GGUF using GGUF format.
Demonstrates building a local financial RAG using Qwen 3.8 27B, Ollama, Qdrant, LangChain and Chainlit on SEC 10-K filings with hybrid search, agentic filtering and streaming.
Locally installing Qwen3.8-27B Obliterated demonstrates risks associated with deploying this uncensored model in production environments.
Speed benchmarks for Ornith 1.5 9B dense and 35B MoE models running on a MacBook M4 Pro with 24 GB unified memory.
Covers Tripo P2.0, mRNA cancer vaccine trial results, Qwen3.8 open weights, OpenAI cyber pacing pause, ChatGPT features, Meta AI app, Perplexity Brain, Gemini updates, Claude integrations, HappyShrimp
Muse Glimmer and Qwen 3.8 27B implement two new features in SuperAsteroids, scored by code review with merging or deleting branches based on results.
Tests experimental Apple Neural Engine Prefill for Qwen3.8-27B on Apple Silicon via oMLX, comparing performance with ANE enabled versus disabled during the prefill stage.
Tests Ox Alpha via OpenRouter using motorcycle simulation, watch website creation, 3D CAD modeling, creative frontend design, ship combat simulation, and multimodal coding tasks.
Thorough testing of DeepSeek-V4-Flash-Vision-Exp.
Tests Ornith-1.5-35B-A3B-GGUF performance, memory, agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot.
Covers Ox Alpha stealth model, GLM-5.3 Flash, Tencent Hunyuan HY4, delayed GPT-6 Astra, held Fable 5.1, Codex limits, Anthropic retention changes, and OpenAI Mona Lisa image model.
Evaluates optimal KV cache quantization using Bob-bench on Mac Studio and 4090 hardware across varying context lengths, referencing a Meta paper.
Installs DeepSeek Harness on Linux, wiring it to Ollama and Unsloth servers running Qwen 3.8-27B at 4-bit locally, then builds a to-do list app in 30 minutes at 16 tokens per second.
Installs and tests Escha-W2, a 2-bit quantized build of Qwen3.8-27B.
Forces BearQ's Explorer Agent to instantly detect new EAApp search features, generate tests via permutations, and update the Application Model using QA Lead and Tester agents.
Evaluates Ornith-1.5-397B-MLX-Q9 performance using Inferencer App on M3 Ultra 512 GiB against claims of surpassing GLM 5.2, Opus and Qwen 3.8 MAX.
Tests dots3-note prev, identified as the first open-weight model in the dots3 family.
Tests Qwen 3.8 27B Q4_K_M reliability via clock tasks, Pomodoro timers, and building a Feed Aggregator in Nim using Ollama and Pi agent.
Tests Qwen3.8-27B local inference on Apple M5 Max, measuring TPS across reasoning levels and evaluating DFlash2 speculative decoding acceleration combined with varying thinking modes.
Locally installs and tests the Ornith-1.5-9B model using resources from huggingface.co/ornith-ai/Ornith-1.5-9B.
Demonstrates building ChatPromptTemplate with runtime placeholders, handling missing variables and JSON curly brace traps via double escaping, partial fills with .partial(), and integrating chat histo
Tests using DeepSeek Harness to call Claude Code and Pi as sub-agents, covering plugin architecture, modes, and custom plugin creation.
Locally installs and tests Ornith-1.5, a model built through end-to-end self-improvement.
Thorough testing of GLM 5.3 covers coding capabilities, security aspects, and creative coding applications using z.ai resources.
Explains QWEN3.8 performance causes, benchmarks it live on M5 Max, and demonstrates replication steps for identical results on local setups.
Tests how Qwen 3.8 27B reasoning levels impact output quality and token usage across tasks including stopwatch, typing speed, sand physics, dungeon crawler, notes app, blender windmill, and charts.
Tests Qwen 3.8 27B against Opus 4.6 on a live game development task using an existing codebase hosted at brrnout.com.
Evaluates Qwen 3.8-27B reasoning_effort levels via Tetris clone, volcano sim, spreadsheet, and drawing tests using identical official thinking samplers on a 17.9 GB 4-bit file.