BBuf/AI-Infra-Auto-Driven-SKILLS
BBuf/AI-Infra-Auto-Driven-SKILLS · 1 plugin
Marketplace Agent-ready skills for LLM serving benchmarks, profiler triage, capacity planning, model Day-0 support, code review, incident triage, and model PR history across SGLang, vLLM, TensorRT-LLM, and TokenSpeed.
Install
The repo has no one-line install. Follow its README.
Plugins 1
After adding the marketplace, install one with /plugin install <name>@ai-infra-auto-driven-skills.
- 1ai-infra-auto-driven-skillsLLM serving benchmarks, SGLang model Day-0 support, profiling, code review, prod incident triage, and PR-history dossiers.
/plugin install ai-infra-auto-driven-skills@ai-infra-auto-driven-skills
Files
AI-Infra-Auto-Driven-SKILLS
Agent-ready skills for LLM serving benchmarks, profiler triage, capacity
planning, model Day-0 support, code review, incident triage, and model PR
history across SGLang, vLLM, TensorRT-LLM, and TokenSpeed.
Plain SKILL.md directories that give a coding agent the operational memory
for real AI-infra work: fair cross-framework benchmarks, kernel-level profiler
reads, operator FLOPs, model support plans, and the upstream PRs that already
solved a similar problem. Kernel campaigns live in the sibling
KDA-Pilot; per-model diffusion runs
live in sglang-diffusion-optimization-flows/.
Releases
v0.1.5
— Release notes (中文), covering model Day-0 support,
profiler layer guides, 118 bilingual model histories, and the October source audit.
Previous release: v0.1.0.
Skills
| Skill | Use it when |
|---|---|
llm-serving-auto-benchmark |
Find the best deployment command for one model across SGLang, vLLM, TensorRT-LLM, TokenSpeed under the same workload, GPUs, and SLA. |
llm-serving-capacity-planner |
Explain startup memory, KV cache budget, request capacity, or OOM pressure from SGLang/vLLM logs. |
llm-torch-profiler-analysis |
Capture or read a torch profiler trace and get kernel, overlap-opportunity, and fusion-opportunity tables checked against a catalog of known SGLang/vLLM/TensorRT-LLM/TokenSpeed/FlashInfer optimizations. |
llm-pipeline-analysis |
Break a trace into forward passes, layers, and kernels with anchor boundaries and Perfetto ranges. |
torch-profiler-layer-track |
Add verified layer-number guides and compact GPU lanes to a trace for Perfetto navigation. |
model-compute-simulation |
Estimate operator shapes, FLOPs, and MFU for a serving shape, or map kernels back to operators. |
sglang-model-day0-support |
Turn a new model architecture into an SGLang Day-0 PR DAG, validation matrix, and release lock. |
sglang-humanize-review |
Review an SGLang PR the way maintainers do, grounded in the full human review corpus. |
sglang-prod-incident-triage |
Turn queue growth, timeouts, wrong outputs, crashes, or stalls into a replay and the next debug step. |
model-architecture-diagram |
Return original public architecture diagrams for popular LLM, VLM, MoE, OCR, and diffusion families. |
Model PR History
model-pr-optimization-history/ is one
queryable knowledge base (installed as model-pr-history-knowledge) with
118 bilingual dossiers: SGLang 45, vLLM
44, TensorRT-LLM 15, TokenSpeed 14.
Each lists a model family's implementation files, every PR that changed them,
and per-PR evidence cards. New generated entries are explicitly marked as
source inventories pending manual diff review. Read it before
patching a model path or calling an optimization new.
cd model-pr-optimization-history
python3 scripts/query.py --list
python3 scripts/query.py --framework vllm "qwen3 fused qk norm"Dossiers are regenerated from upstream git history with
tools/rebuild_model_pr_history_from_git.py; see
update_prompt.md for the full refresh procedure and
docs/upstream-source-contracts.md for
the inspected source revisions.
Install
Claude Code plugin:
/plugin marketplace add BBuf/AI-Infra-Auto-Driven-SKILLS
/plugin install ai-infra-auto-driven-skills@ai-infra-auto-driven-skills
Any skill runtime (Claude Code, Codex, Kimi, ...): link or copy the skill
directories into its skill directory.
git clone https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git
cd AI-Infra-Auto-Driven-SKILLS
SKILL_DIR=~/.claude/skills # or ${CODEX_HOME:-~/.codex}/skills
mkdir -p "$SKILL_DIR"
for d in skills/*/ skills/model-optimization/*/; do
[ -f "$d/SKILL.md" ] && ln -sfn "$PWD/${d%/}" "$SKILL_DIR/$(basename "$d")"
done
ln -sfn "$PWD/model-pr-optimization-history" "$SKILL_DIR/model-pr-history-knowledge"Evidence Rules
- Benchmark rows record model, framework commit, GPUs, workload, rate or
concurrency, SLA status, both commands, and raw artifacts. - Profiler reports keep prefill and decode separate and never reuse an older
trace for a new capture. - Performance claims are scoped to the exact model, hardware, precision,
workload, and framework revisions; accuracy is checked on the real path. - Historical PR evidence keeps its audit date; a source refresh is not a GPU
rerun.
Related Projects
- KDA-Pilot hosts standalone kernel
loops, kernel knowledge, and NCU workflows.
Star History
{
"name": "ai-infra-auto-driven-skills",
"description": "Agent-ready LLM serving, SGLang model Day-0 support, profiling, capacity, incident triage, architecture, and PR-history skills.",
"owner": {
"name": "BBuf",
"url": "https://github.com/BBuf"
},
"plugins": [
{
"name": "ai-infra-auto-driven-skills",
"source": "./",
"description": "LLM serving benchmarks, SGLang model Day-0 support, profiling, code review, prod incident triage, and PR-history dossiers.",
"version": "0.8.0",
"category": "ai-infra",
"tags": [
"sglang",
"day0",
"vllm",
"tensorrt-llm",
"benchmarks",
"profiler",
"kernel",
"moe",
"llm-serving"
]
}
]
}Facts
- Kind
- Marketplace
- Repo
- BBuf/AI-Infra-Auto-Driven-SKILLS
- Group
- Uncategorized
- Marketplace name
- ai-infra-auto-driven-skills
- Owner
- BBuf
- Language
- Python
- Created
- 2026-04-01
- Forks
- 83
- Plugins
- 1
- 1f/prompts.chatf/prompts.chatf.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
- 2affaan-m/everything-claude-codeaffaan-m/everything-claude-codeThe agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
- 3obra/superpowersobra/superpowersAn agentic skills framework & software development methodology that works.
- 4anthropics/skillsanthropics/skillsPublic repository for Agent Skills
- 5anthropics/claude-codeanthropics/claude-codeClaude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
- 6nextlevelbuilder/ui-ux-pro-max-skillnextlevelbuilder/ui-ux-pro-max-skillAn AI skill that provides design intelligence for building professional UI/UX across multiple platforms.