Dirk Qwen 3.8 27B tested - Local LLM setup
Tests Dirk Qwen 3.8 27B GGUF from peculiar-ragdoll using Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot benchmarks.
Videos
Find videos about LLM models, hardware, runtimes, benchmarks and practical AI workflows.
Tests Dirk Qwen 3.8 27B GGUF from peculiar-ragdoll using Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot benchmarks.
Builds Pydantic models with with_structured_output using Nemotron 3.5 Lightning on OpenRouter to generate consistent JSON fields for sentiment analysis and ticket routing via LangChain.
Locally installs and tests Thomson-1.0-Small, a model designed for high-stakes professional work across legal, tax, and journalism domains.
An AI coding agent uses Cursor, Node.js 24, TypeScript and Zapier TypeScript SDK to move a Google Calendar event and DM attendees via Slack.
Demonstrates setting up Pithagoras, an open-source local AI agent with native browser control, using four onboarding components: identity, intake, action access, and escalation path.
Locally installs and tests Pipecat's PhoneLLM Alpha 1, an open-weights model for voice agent use cases.
Compares DeepSeek V4 Flash and Qwen 3.8 Flash Next using Unsloth UD-IQ1_S 1-bit builds on 19 hard problems via CPU-only and RTX 5090 expert layer offload benchmarks.
Hands-on guide to training kyutai pockettts models from scratch for custom languages and voices using local resources.
Covers GLM 5.3 Flash, Qwen 3.8 Flash Next, Minimax FastH3, Hy4, Ox Alpha, Gemini 3.5 Transcribe, Omni 1.1 Flash, Block 3D, FixAnything, and World Humanoid Games.
Benchmark IBM Granite 4.2, Qwen 3.8 27B, Gemma 4, and Ornith 1.5 on an RTX 5090, measuring KV cache cost, usable context, generation speed, coding, math, and tool calling via Python tests.
Hands-on comparison of GLM-5.3-Flash vs Qwen3.8-Flash evaluating bug fixing, vision coding capabilities, and AI flirting performance through three distinct tests.
Demonstrates LCEL and the pipe operator to build sequential and parallel chains using StrOutputParser, .batch(), .stream(), and model fallbacks, verified via LangSmith.
Runs GLM-5-Flash using Inferencer App v2.3.5 on an M3 Ultra 512 GiB system with the inferencerlabs/GLM-5.3-Flash-MLX-Q9 quantization via local testing.
Tests Hy4 preview, a new-generation Mixture-of-Experts (MoE) flagship model developed by the Tencent Hy Team.
Runs Qwen3.8 27B locally using Superlinked Inference Engine (SIE) for efficient execution.
Compares MLX/oMLX and GGUF with llama.cpp for local inference of Qwen3.8-Flash-Next 125B on Apple M5 Max, achieving ~60 tokens/sec with oMLX 0.6.3 and Lightning MTP.
Tests Qwen 3.8 27B Ridge from Empero AI on 16GB GPUs using benchmarks including OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot.
Benchmarking Ornith 1.5 9B and 35B, Qwen 3.8 27B, and Gemma 4 on an RTX 5090 using llama-server with 4-bit weights and FP16 KV cache via code-graded metrics.
Tests Qwen 3.8-Flash-Next 1-bit to 4-bit IQ4_XS via llama.cpp on GMKtec EVO-X2 against Qwen 3.8-27B references using four one-shot tests, reporting ~20 tok/s and peak memory usage.
Step-by-step guide to running unsloth/GLM-5.3-Flash-GGUF locally using CPU and RAM.
Tests GLM5.3 Flash performance on browser workflows, game generation, C++ coding, model training, and Blender & Godot Street Yeet tasks.
Benchmarks Qwen 3.8 Flash Next on Intel Core Ultra 9 285K CPU versus RTX 5090 offloading using llama.cpp, measuring throughput and perplexity for 1-bit and 2-bit quantizations.
Compares Superlinked Inference Engine (SIE) with vLLM.
Evaluates Qwen3.8-Flash-Next-MLX-Q9 and Q4 variants using Inferencer App v2.3.4 on an M3 Ultra 512 GiB system via local testing.
Weight-level teardown shows 131 of 866 tensors changed in Qwen 3.8 27B Uncensored via abliteration, tested on RTX 5090 with llama-server against a 30-prompt refusal benchmark.
Tests Qwen3.8-Flash-Next coding and Splunk Cyber CTF performance, investigates hallucinations caused by Qwen Sparse Attention, and demonstrates fixing issues via sparse attention settings.
Reviews Qwen 3.8-Flash-Next specs and Unsloth GGUF sizes.
Demonstrates building coding agent hooks via the hooks-create skill to enforce workflow guarantees through events, scripts, and wiring for tests and security.
Locally tests GLM-5.3-Flash, identified as the first natively multimodal model in the GLM-5 series.
Tests Qwen3.8 Flash Next on browser workflows, C++ games, Blender, Godot, local Q4, FPS generation, and 3D scenes.
Locally installs and tests Qwen3.8-Flash-Next.
Explores Qwen3.8-Flash-Next architecture and runs local inference via Unsloth GGUF and llama.cpp on Apple Silicon, covering 1-bit and 4-bit quantization setups.
Compares Qwen 3.8 Flash Next and Qwen 3.8 27B architectures, detailing MoE, sparse attention, gated residuals, 51B n-gram lookup tables, and RTX 5090 MTP measurements.
Tests combining multiple GPUs into a single desktop setup with 224GB of memory to evaluate performance beyond intended specifications using Strix Halo, LLM and Nvidia components.
Integrates Model Context Protocol into the TalkWithMe Star Trek group chat using mcp-light to enable AI personas to perform actions within an agentic loop rather than passive chatting.
Explores running a 26B parameter MoE model on limited memory, comparing weight sizes across FP16/BF16, 8-bit, 4-bit, 3-bit, and 2-bit quantization levels.
Demonstrates running Qwen 3.8 27B on consumer-grade computers using LM Studio to select compatible high-quality quantizations and manage downloads.
Evaluates Qwen3.8 27B Cold Fusion GAIN V1.1 NM DAU NEO MAX MTP against the base model using Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender, and Godot.
25:06
Evaluates GLM 5.3, Qwen 3.8 Max and Grok 4.6 using ten hard briefs in Blender, Unity and Godot to assess reliability and capability in building real 3D projects.
Hands-on guide demonstrating local fine-tuning of the Qwen3.8 27B model using custom datasets.
Evaluates Qwen 3.8-27B MTP depths 2, 3, 4 versus off using llama.cpp on GMKtec EVO-X2, measuring throughput and verifying output identity via hashes under greedy and sampled settings.
Compares Qwen 3.8, 3.6 and 3.5 using 397 generations on an RTX 5090, analyzing MTP layer activation, split vision encoder, post-training effects, Terminal-Bench scores and token speeds.
Locally installs and tests Granite 4.2 generation in 3B and 8B variants using ibm-granite/granite-4.2-3b resources.
Demonstrates cleaning reference audio with UltimateVocalRemover and integrating OmniVoice into the TalkWithMe app for offline, real-time conversations with AI personas using cloned voices.
Tests Qwen 3.8 27B, Muse Glimmer 30B and Gemma 4 26B on an RTX 5090 using Ollama, evaluating consistency across 12 accounting questions run 10 times each.
Locally installs and tests FreeToken, an edge-native Mixture-of-Experts (MoE) serving engine.
Evaluates ClinePass for accessing open-weight coding models within Cline, testing performance on a real-world project using Plan and Act modes alongside the World of AI Benchmark tool.
Locally installs and tests EVIE-Preview-4.5B, a visual document retriever featuring native 128-dimensional token vectors.
Evaluates whether the Ornith 1.5 35B A3B mixture of experts model surpasses the original 1.0 version through a coding challenge comparison.
Performs unsloth/Qwen3.8-27B-GGUF quantizations Q1-Q8 using HumanEval within Kanban, Blender and Godot environments.