Bring any model
Use DeepSeek, OpenAI, GLM, Kimi, or an OpenAI-compatible endpoint. CoDebugger adds the debugging system around it.
CODE AGENT FOR AI INFRA ENGINEERS
Trace production failures across code, Kubernetes, logs, schedulers, distributed runtimes, and accelerators with the model you already trust.
parallel_config.py:184
WHY ANOTHER CODE AGENT?
Use DeepSeek, OpenAI, GLM, Kimi, or an OpenAI-compatible endpoint. CoDebugger adds the debugging system around it.
Correlate selected repositories with read-only Kubernetes state, component logs, CRDs, events, and exact code paths.
Skip generic codebase exploration. Architecture maps and validation rules direct the agent to the decisive components first.
One failed GPU hour is expensive.
FIELD-TESTED KNOWLEDGE
CoDebugger’s AI Infra knowledge is curated from real failures handled by frontline engineers: scheduler predicates, distributed topology constraints, operator reconciliation, NCCL/HCCL communication, inference engine startup, and accelerator runtime failure modes.
A CONTROLLED DEBUG LOOP
Choose one or more local repositories and the exact Git branches.
Optionally connect a cluster with an enforced read-only Kubernetes toolset.
The agent searches code, reads files and logs, and validates configuration assumptions.
Get a structured root cause, confirmed evidence, code locations, and next checks.
STOP DEBUGGING AI INFRA FROM SCRATCH