Loading...
Loading...
Load for any work involving Baseten - deploying/operating models on Dedicated Inference (Truss, custom Docker servers, TRT-LLM engines, Chains), calling hosted Model APIs, running Training jobs (SFT/RL/LoRA), or Model Frontier Gateway.
npx skill4agent add basetenlabs/baseten-skills baseten| Component | Provides | Install |
|---|---|---|
| Interact with backend (~REST API, CRUD): models, deployments, training, environments, secrets, chains. API-key auth. | |
| Semantic search + filesystem of | |
| Needed for model/chain push from local code, watch (= live patch). Needs | |
| | reachable via HTTP |
| This skill | | |
BASETEN_MCP_KEYapp.baseten.co/settings/api_keystruss --version--remote <name>truss [subcommand] --helpllms.txttruss watch| You want… | Flavor | Specialization | When NOT to pick |
|---|---|---|---|
| Hosted LLM, no deploy step | Model APIs | | model not in catalog; need custom hardware, requirements, stability... |
| LLM on an off-the-shelf server (vLLM / SGLang / TGI / Triton / NIM) | Custom Docker server | | the server doesn't exist or you need Python in the request path |
| LLM/embedding via a Baseten engine (TRT-LLM / BEI / BIS-LLM); minimal config, no Python | Engine-only | | architecture not covered by an engine; you need custom logic |
| Custom Python in the request path (pre/post, custom arch, weird IO) | Python-class Truss ( | | an engine or off-the-shelf server fits — pick that, it's faster to ship |
| Multi-step pipeline with heterogeneous hardware / per-step scaling (RAG, ASR→LLM→TTS, fan-out, chunking) | Chains | | one-stage or homogeneous — a single Truss is simpler |
model-dev-loop.mddeployment-lifecycle.mddeployment/rolling-deployments.mdxinference-api.mdmodel-apis.mdinference/performance-client.mdxmanagement-api.mdmodel.pyls references/references/truss-cli.mdtruss pushwatchreferences/truss-config.mdreferences/truss-model-py.mdreferences/truss-custom-servers.mddocker_serverreferences/truss-chains.mdreferences/model-apis.mdreferences/inference-api.mdrequestshttpxreferences/management-api.mdtrussreferences/deployment-lifecycle.mdreferences/model-dev-loop.mdPulling imagemodel_cache: Fetch tookCompleted model.load() execution in N ms| Source | Strength | Gap / quirk |
|---|---|---|
| API specs, protocol details, knobs | Lags product; no perf numbers |
| Flagship managed models, perf claims | Some entries are sales-gated, not self-serve |
| What's actually one-click API-deployable | Doesn't include every marketing-library entry; lacking tags |
| Concrete latency / cost / vs-competitor numbers, technical deep dives | Unstructured; not in docs MCP. Discover via |
| High-level pitch | May describe flagship features that need a Baseten engagement |
| Working code patterns | Often outdated / broken / drifted. Consult with caution, last resort. Might need fixups before deploy works. |
list_library_modelsbaseten.co/llms.txtdocs.baseten.co/llms.txtreferences/truss-config.mdmodel_cachetruss-config.mdtruss-chains.mdbasetenbaseten_docstraining/overview.mdxloops/overview.mdxtraining/ssh.mdxtraining/remote-access.mdxbaseten_docsdocs.baseten.co/llms.txtcathead.mdxquery_docs_filesystem_basetendocs.baseten.co/<path>.mdmodel-dev-loop.mdlist_library_modelsdisplay_namehf_repo_idbaseten.co/llms.txtget_deploymentget_deployment_logsget_deployment_configlist_*curl