gpu-mcp-server

MCP serverDocs & knowledge

This app lets your AI read live metrics from NVIDIA GPUs: utilization, memory use, temperature, and power draw. It also supports MIG, where one physical GPU is split into smaller isolated instances. Once added, you can ask your AI how your GPUs are doing and get the numbers directly.

Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.

After adding it, ask your AI for the GPU stats you care about, such as utilization, memory, temperature, or power on a specific card.

What your AI can do with it

  • Report current utilization for NVIDIA GPUs
  • Check how much GPU memory is in use
  • Read GPU temperatures
  • Show power draw for each GPU
  • View metrics for MIG-partitioned GPUs

From the project's README

As published by pmady/gpu-mcp-server in README.md.

Note: the OpenSSF Best Practices questionnaire is in progress. Once the project entry is registered at https://www.bestpractices.dev/en, swap the static badge above for the live one: [![OpenSSF Best Practices](https://www.bestpractices.dev/projects/<ID>/badge)](https://www.bestpractices.dev/projects/<ID>)

An MCP server that exposes NVIDIA GPU metrics as tools. Any MCP-compatible AI agent (Claude, Goose, Cursor, etc.) can query real-time GPU utilization, memory, temperature, power, PCIe and NVLink throughput no Prometheus or dcgm-exporter required.

Built on the official Go MCP SDK and NVIDIA go-nvml.

Tools

ToolDescription
list_gpusList all GPUs with utilization and memory info
get_gpu_metricsDetailed metrics for a GPU by index or UUID
get_gpu_processesPID-level GPU process attribution
gpu_summaryAggregate stats across all devices

All tools support MIG (Multi-Instance GPU) - MIG instances appear as separate devices with their parent GPU's shared metrics (temperature, power, PCIe).

Sample output

Each tool returns structured JSON. The examples below show the shape of the data an agent receives from a node with two NVIDIA A100 GPUs.

list_gpus:

{
  "count": 2,
  "devices": [
    {
      "index": 0,
      "uuid": "GPU-aaaa-1111",
      "name": "NVIDIA A100-SXM4-80GB",
      "gpu_utilization_percent": 85,
      "memory_used_mib": 57344,
      "memory_total_mib": 81920
    },
    {
      "index": 1,
      "uuid": "GPU-bbbb-2222",
      "name": "NVIDIA A100-SXM4-80GB",
      "gpu_utilization_percent": 20,
      "memory_used_mib": 12288,
      "memory_total_mib": 81920
    }
  ]
}

get_gpu_metrics (with {"index": 0} or {"uuid": "GPU-aaaa-1111"}):

{
  "index": 0,
  "uuid": "GPU-aaaa-1111",
  "name": "NVIDIA A100-SXM4-80GB",
  "gpu_utilization_percent": 85,
  "memory_utilization_percent": 70,
  "memory_used_mib": 57344,
  "memory_total_mib": 81920,
  "temperature_celsius": 72,
  "power_draw_watts": 300,
  "power_limit_watts": 400,
  "pcie_tx_kbps": 0,
  "pcie_rx_kbps": 0,
  "nvlink_tx_mbps": 0,
  "nvlink_rx_mbps": 0
}

gpu_summary:

{
  "device_count": 2,
  "avg_gpu_utilization": 52.5,
  "avg_memory_utilization": 42.5,
  "total_memory_used_mib": 69632,
  "total_memory_total_mib": 163840,
  "max_temperature_celsius": 72,
  "total_power_draw_watts": 375
}

MIG instances add is_mig, parent_gpu, and mig_profile fields to the get_gpu_metrics and list_gpus payloads.

Quick start

# build (requires CGO + NVML headers on Linux)
make build

# run the server communicates over stdio
./gpu-mcp-server

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "gpu": {
      "command": "/path/to/gpu-mcp-server"
    }
  }
}

Goose

extensions:
  gpu-metrics:
    type: stdio
    cmd: /path/to/gpu-mcp-server

Cursor

Add to .cursor/mcp.json for a project, or ~/.cursor/mcp.json for all projects:

{
  "mcpServers": {
    "gpu": {
      "type": "stdio",
      "command": "/path/to/gpu-mcp-server"
    }
  }
}

Windsurf

Add to ~/.codeium/windsurf/mcp_config.json:

{
  "mcpServers": {
    "gpu": {
      "command": "/path/to/gpu-mcp-server"
    }
  }
}

Build

Requires Go 1.23+, CGO, and NVIDIA drivers on the target machine.

make build       # compile binary
make test        # run tests (no GPU needed uses mock)
make lint        # golangci-lint
make docker      # container image

Tests use a mock collector, so they run anywhere no GPU hardware required.

Docker

Prebuilt multi-arch images (linux/amd64, linux/arm64) are published to GHCR on every release.

docker pull ghcr.io/pmady/gpu-mcp-server:latest
docker run --rm -i --gpus all ghcr.io/pmady/gpu-mcp-server:latest

The host needs the NVIDIA Container Toolkit installed for --gpus all to work. The server speaks MCP over stdio, so the -i flag is required — don't drop it.

{
  "mcpServers": {
    "gpu": {
      "command": "docker",
      "args": ["run", "--rm", "-i", "--gpus", "all", "ghcr.io/pmady/gpu-mcp-server:latest"]
    }
  }
}

Pin a specific version via tag instead of :latest, e.g. ghcr.io/pmady/gpu-mcp-server:v0.1.0.

Architecture

Agent (Claude/Goose) ─── MCP (stdio) ──→ gpu-mcp-server ──→ NVML ──→ GPU
                                              │
                                         Tools:
                                         • list_gpus
                                         • get_gpu_metrics
                                         • gpu_summary

The server runs as a local process alongside the agent. It calls NVML directly through cgo — no sidecar, no network hops, no metric pipeline to configure.

Project info

Roadmap

See ROADMAP.md for the 12-month public roadmap.

Contributing

See CONTRIBUTING.md for how to get involved.

Contributors

Thanks to all our contributors! Add yourself via PR.

Governance

This project follows Linux Foundation Minimum Viable Governance.

Documentation

Star History

Signals

GitHub stars
15
Forks
13
Last commit
Aug 2026
Advanced
Delivery
gpu-mcp-server MCP server → your ahel gateway (mcp.ahel.ai) → every connected AI client.
Catalog kind
mcp-server
Gateway key
io-github-pmady-gpu-mcp-server
Source
github.com/pmady/gpu-mcp-server