🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
-
Updated
Oct 7, 2026 - Python
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
macOS menubar app for fast local DeepSeek V4.1, with 1M context.
⚡️ A community driven PHP client for DeepSeek AI, designed to bring clean API access, fluent developer experience, and framework-friendly integration to PHP applications.
Fixes missing reasoning_content for DeepSeek V4
基于 FastAPI 的 DeepSeek Chat 反向代理,将 DeepSeek 网页版的 API 转换为 OpenAI 兼容格式。 支持流式/非流式对话、专家模式、深度思考(reasoning_content)、工具调用(DSML prompt injection)。 自动处理 PoW 鉴权挑战,无需官方 API Key。
DeepSeek-V4-Flash-0731 284B inference in ~25 GB of RAM / Qwen3.8-Next-Flash-FP8 inference in ~18 GB of RAM on any M-series MacBook
DeepSeek V4 Flash CPU/NVMe research fork: 78.62 GiB GGUF validated on 7.7 GiB RAM, CPU-only, using demand paging.
My public AI learning journey — sharing what I learn, build, test, and discover across AI. Traffic: 7K+ monthly clicks · 150K+ AI citations/month · 5.5 average search position across search engines. local-ai-zone.github.io
A tool to have multiple claude-code instance with deepseek, minimax, and z.ai glm models
Winning Qwen and DeepSeek setups for AMD Strix Halo: pinned Ansible recipes and matched benchmarks on a 128 GiB Ryzen AI Max+ 395 / Radeon 8060S.
DeepSeek 逆向 API 支持 Deepseek V4
Guide for development with DeepSeek Harness. Building plugin for DeepSeek Harness Project.
Reproducible kit to deploy DeepSeek-V4-Flash-DSpark on a 2× NVIDIA DGX Spark (GB10) cluster: vLLM TP=2 over QSFP 200GbE, NVFP4 KV, DSpark speculative decoding, 1M context, systemd self-heal. Apache-2.0.
A collection of recipes/notebooks showcasing use-cases of open-source models with Qubrid AI.
Codex 桌面版为 DeepSeek-V4-Flash 开启 Max 推理档位的完整排障与配置指南 / Complete guide to enable Max reasoning effort for DeepSeek-V4-Flash in Codex desktop
Codex vision bridge for DeepSeek V4 Flash: give text-only DeepSeek image capability in Codex. Local proxy turns pasted images and view_image into text via free GLM-4V-Flash or any OpenAI-compatible vision API. No GPU, no Ollama.
Dynamic Agent-to-Agent (A2A) task graph generation, subtask independence verification, and parallel multi-agent orchestration
DeepSeek-V4-Flash-0731 on 8x RTX 3090 (SM86) with vLLM — verified FP8 serving, benchmarks, build guide, and reproducible release.
To associate your repository with the deepseek-v4-flash topic, visit your repo's landing page and select "manage topics."