← Back to skills

CURATED FROM PUBLIC GIT

doca-gpunetio-ib-write-bw

>

↓ 0★ 0by greatsage_sh
893c384d4c03
skillmarket install greatsage_sh/doca-gpunetio-ib-write-bw
Public Git sourcehttps://github.com/NVIDIA/skills.gitcommit 893c384d4c03770e7c13b4dcee1de6cd16e7a9c0
Part of skillsetskills →

About this skill

---
license: Apache-2.0
name: doca-gpunetio-ib-write-bw
description: >
  Use this skill when the user is building, running, or interpreting
  the doca/tools/gpunetio_ib_write_bw client+server benchmark — a CUDA
  kernel on the server posts RDMA WRITE work requests through the
  doca-gpunetio device-side surface to measure sustained GPU-driven
  WRITE bandwidth on a GPU+IB-device pair. Trigger even when the user
  does not explicitly mention "doca-gpunetio-ib-write-bw" or
  "GPUNetIO" — typical implicit phrasings include "measure WRITE BW
  when the GPU posts the WRs", "BW swings between runs on the same
  flags", "is the NIC saturated or am I CPU-bound on the CUDA
  kernel", "meson compile fails for the GPUNetIO bw tool",
  "nvidia_peermem isn't picking up my GPU buffer", or "GPU-initiated
  WRITE throughput vs CPU-initiated perftest". Refuse and route
  elsewhere for general doca-gpunetio library work, DOCA install, the
  GPU-initiated WRITE latency analog, the CPU-initiated upstream
  perftest, or application-level end-to-end throughput — those belong
  to other skills.
metadata:
  kind: tool
compatibility: >
  Requires DOCA SDK installed at /opt/mellanox/doca on Linux (Ubuntu
  22.04/24.04 or RHEL/SLES) with a BlueField DPU or ConnectX NIC. Reads
  `pkg-config doca-gpunetio doca-rdma doca-common` and inspects
  /opt/mellanox/doca/tools/gpunetio_ib_write_bw/. Requires an NVIDIA
  GPU with CUDA toolkit + nvcc installed, the `nvidia_peermem` kernel
  module loaded, and an InfiniBand-capable RNIC paired with the GPU on
  a common PCIe/NVLink fabric.
---

# DOCA GPUNetIO ib_write_bw

**Where to start:** This is a tool skill for the GPUNetIO-
flavored `ib_write_bw` benchmark shipped under
`doca/tools/gpunetio_ib_write_bw/` (a client + server pair,
built from source against the installed DOCA via `meson`).
It measures sustained RDMA WRITE bandwidth when the WRs are
posted **from a CUDA kernel through the doca-gpunetio
device-side surface**, with the GPU on the data path. Open
[`TASKS.md`](TASKS.md) and start at
[`## configure`](TASKS.md#configure) for the GPU-NIC
pairing precondition and the build pattern; jump to
[`## run`](TASKS.md#run) for the smoke-before-bulk flow.
Open [`CAPABILITIES.md`](CAPABILITIES.md) when the question
is *what this tool actually measures*, *how the result
decomposes (GPU occupancy vs NIC issue rate vs link
saturation)*, or *how the result reads against the GPI
sister tool and the upstream CPU-initiated `perftest`
`ib_write_bw`*. If DOCA is not installed yet, route to
[`doca-setup`](../../doca-setup/SKILL.md) first; if the
user is still deciding between the GPI and GPUNetIO
programming surfaces, the picture in
[`../../libs/doca-gpunetio/CAPABILITIES.md#capabilities-and-modes`](../../libs/doca-gpunetio/CAPABILITIES.md#capabilities-and-modes)
and
[`../../libs/doca-gpi/CAPABILITIES.md#capabilities-and-modes`](../../libs/doca-gpi/CAPABILITIES.md#capabilities-and-modes)
is the first stop.

## Example questions this skill answers well

The CLASSES of `doca-gpunetio-ib-write-bw` questions this
skill is built to answer, each with one worked example. The
class is the load-bearing piece; the worked example is one
instance.

- **"What sustained RDMA-WRITE bandwidth can the GPUNetIO
  path deliver on this GPU-NIC pair?"** — worked example:
  *"measure sustained WRITE BW between two hosts with an
  H100 + ConnectX-7 on each side"*. Answered by the
  GPU-NIC pairing precondition in
  [`CAPABILITIES.md ## Capabilities and modes`](CAPABILITIES.md#capabilities-and-modes)
  + the bring-up flow in
  [`TASKS.md ## configure`](TASKS.md#configure) +
  [`TASKS.md ## run`](TASKS.md#run). The same shape
  answers *"measure GPUNetIO-driven WRITE BW between a
  host GPU and a BlueField DPU"*.
- **"Where is the bottleneck — GPU compute occupancy, NIC
  issue rate, or link saturation?"** — worked example:
  *"I see 120 Gbit/s on a 200 Gbit/s link; is the NIC
  saturated, am I CPU-bound on the client, or is the CUDA
  kernel not driving enough WRs in flight?"*. Answered by
  the throughput-decomposition rules in
  [`CAPABILITIES.md ## Observability`](CAPABILITIES.md#observability)
  + the eval-loop overlay in
  [`TASKS.md ## test`](TASKS.md#test).
- **"How does the result differ from the classic CPU-
  initiated `perftest` `ib_write_bw`?"** — worked example:
  *"my team has a CPU-initiated WRITE BW number on this
  same NIC; should I expect the GPUNetIO number to match
  or be different?"*. Answered by the *"GPU-initiated
  path adds (or removes) overhead vs the CPU-initiated
  path"* rule in
  [`CAPABILITIES.md ## Capabilities and modes`](CAPABILITIES.md#capabilities-and-modes).
- **"Is the doca-gpunetio path the right surface for my
  sustained-throughput workload class?"** — worked example:
  *"my application streams sensor data from GPU memory at
  line rate to a remote consumer"*. Answered by the
  *"when GPUNetIO is the right surface vs GPI vs CPU-
  initiated"* rule in
  [`CAPABILITIES.md ## Capabilities and modes`](CAPABILITIES.md#capabilities-and-modes)
  + the use-side decision in [`TASKS.md ## use`](TASKS.md#use).
- **"My BW number swings between runs. What do I check
  before quoting it?"** — worked example: *"three runs at
  the same flags gave 145, 187, and 160 Gbit/s; is the
  benchmark noisy or is my platform inconsistent?"*.
  Answered by the measurement-soundness rules in
  [`CAPABILITIES.md ## Error taxonomy`](CAPABILITIES.md#error-taxonomy)
  layer 5 + the steady-state guidance in
  [`TASKS.md ## test`](TASKS.md#test).
- **"What version of DOCA + CUDA Toolkit do I need for this
  binary to build and run?"** — worked example: *"my
  install has DOCA at one semver and CUDA at another; will
  the ToT-shipped `gpunetio_ib_write_bw` even link?"*.
  Answered by the version overlay in
  [`CAPABILITIES.md ## Version compatibility`](CAPABILITIES.md#version-compatibility)
  which cross-links the canonical detection chain in
  [`doca-version`](../../doca-version/SKILL.md).

## Audience

This skill serves **external developers and performance
engineers who need a reproducible measurement of sustained
RDMA WRITE bandwidth when the WRs are posted from a CUDA
kernel through doca-gpunetio**, on the user's actual install
and GPU-NIC pair. Concretely:

- A developer comparing the GPUNetIO path against the GPI
  path or the host-initiated `perftest`-style path before
  committing an application design to one of them.
- A platform operator validating a tuning change (NUMA
  pinning, GPU PCIe placement, IB device choice, GID
  index, NIC firmware burn) by re-running this benchmark
  against the new state.
- An SRE / performance engineer producing a *"this is the
  GPUNetIO-driven WRITE BW on this GPU-NIC pair today"*
  artifact downstream consumers can cite.
- An AI agent answering *"is the doca-gpunetio path a win
  for my sustained-throughput workload class"* honestly —
  with a measured number, the build + invocation that
  produced it, and the GPU + NIC + DOCA version that
  scopes it — rather than guessing from datasheet
  headlines.

It is **not** for users debugging the `doca-gpunetio`
library itself (route to
[`../../libs/doca-gpunetio/SKILL.md`](../../libs/doca-gpunetio/SKILL.md)),
and **not** a substitute for the `perftest` upstream
`ib_write_bw` (which measures CPU-initiated WRITE BW).

## Language scope

The `doca-gpunetio-ib-write-bw` tool is shipped as **C plus
a CUDA `.cu` translation unit** under
`doca/tools/gpunetio_ib_write_bw/`, split into a `client/`
subtree and a `server/` subtree. The verified surface (per
`client/{main.c,common.h,common.c,kernel.cu,perftest.c}` and
`server/{main.c,common.h,common.c,perftest.c}`): host-side
build via `meson` against the installed DOCA `pkg-config`
modules (`doca-gpunetio`, `doca-rdma`, `doca-common`); the
device-side build via `nvcc` against the DOCA GPU NetIO
device-side header set; the OOB descriptor exchange via a
TCP socket between client and server. There is no Python /
Rust / Go binding — the tool is a pair of CLI binaries.
The skill's job is to keep the operator-side workflow
language-neutral; the device-side CUDA surface is not
wrappable in another language.

## When to load this skill

Load this skill when the user is — or the agent needs to —
build and run the `gpunetio_ib_write_bw` client + server on
real hosts with DOCA installed plus a CUDA Toolkit matched
to the DOCA install, and a GPU + IB device pair on the
host's PCIe topology. Concretely:

- Measuring sustained kernel-initiated RDMA WRITE
  bandwidth between two hosts (or a host and a BlueField
  DPU) with the GPUNetIO surface.
- Deciding whether the GPUNetIO path is the right runtime
  surface for a class of workload vs the GPI programming
  surface (the [`doca-gpi`](../../libs/doca-gpi/SKILL.md)
  library — `doca/tools/` ships no GPI benchmark binary) or
  the classic CPU-initiated `perftest` path.
- Capturing a documented baseline (build + invocation +
  DOCA version + GPU + NIC + as-deployed environment +
  numbers) for later regression hunts.
- Diagnosing a build / link / run failure that surfaces
  the GPUNetIO + RDMA bring-up sequence under this tool's
  shipped scaffolding.

Do **not** load this skill for general DOCA orientation,
library API work, or installation. For those, use
[`doca-public-knowledge-map`](../../doca-public-knowledge-map/SKILL.md),
[`../../libs/doca-gpunetio/SKILL.md`](../../libs/doca-gpunetio/SKILL.md),
or [`doca-setup`](../../doca-setup/SKILL.md). Do not load
it for *application-level* end-to-end throughput either —
this benchmark measures the WR-submission path through
GPUNetIO, not the user's full pipeline.

## What this skill provides

This is a **thin loader**. Substantive material lives in
two companion files:

- `CAPABILITIES.md` — what the tool measures (the
  sustained-WRITE-BW primitive driven by a server-side
  CUDA kernel through doca-gpunetio), the
  runtime-surface selection rule (GPUNetIO vs GPI vs
  CPU-initiated), the GPU-NIC pairing precondition, the
  throughput-decomposition guide (GPU compute occupancy
  vs NIC issue rate vs link saturation), the version
  overlay (DOCA `.pc` PLUS CUDA Toolkit), the layered
  error taxonomy (config-syntax / build-time / GPU-NIC-
  pairing / GPUNetIO-lifecycle / RDMA-connection /
  measurement-soundness / version / cross-cutting), the
  observability surface (stdout report, DOCA log levels,
  OOB-socket exchange), and the safety overlay (the
  *"GPU-side handle is a credential"* rule from
  doca-gpunetio; the cross-cutting hardware-safety
  meta-policy).
- `TASKS.md` — step-by-step workflows for the in-scope
  task verbs: `install` (preconditions — DOCA install,
  CUDA Toolkit, GPU + NIC pair, OOB connectivity),
  `configure` (build-tree under
  `doca/tools/gpunetio_ib_write_bw/` and the `meson`
  build wrapping the shipped DOCA), `build` (the
  `meson setup` + `meson compile` pattern from the
  public DOCA build documentation), `modify` (do not
  patch the shipped tool source; modify the invocation
  and the surrounding environment instead), `run` (smoke-
  before-bulk; client + server bring-up order; reading
  the per-iteration report), `test` (the eval loop —
  steady-state, NUMA placement, NIC saturation cross-
  check), `debug` (walk the error taxonomy layer by
  layer), `use` (how a BW result feeds a class-of-
  workload decision), plus a `Deferred task verbs`
  block routing out-of-scope questions.

The skill assumes a host where DOCA is already installed,
a CUDA Toolkit matched to the install is present, and the
operator has whatever privileges the public install profile
expects for binding a `doca_dev`, a `doca_gpu`, and an OOB
TCP socket.

## What this skill deliberately does not ship

This skill is **agent guidance**, not a samples or scripts
bundle. To keep the boundary clean, it deliberately does
not contain — and pull requests should not add:

- **Specific flag strings or expected throughput numbers**
  beyond what the tool's shipped `--help` and `main.c` ARGP
  registration establish. The flag surface is small
  (device name, GPU PCIe address, GID index, server IP on
  the client side); the agent re-reads the binary's
  `--help` on the installed version before quoting flag
  strings. Throughput numbers are device-, firmware-,
  version-, and topology-specific.
- **Pre-written DOCA GPUNetIO or CUDA kernel source code**
  that would compete with the shipped tool tree. The
  shipped `client/{main.c,kernel.cu,perftest.c,common.{c,h}}`
  and `server/{main.c,perftest.c,common.{c,h}}` files are
  the verified worked example; the agent's job is to
  route the user there and prescribe minimum-diff
  modification per the universal modify-a-sample workflow
  in
  [`doca-programming-guide`](../../doca-programming-guide/SKILL.md).
- **Wrappers, parsers, or scripts** in any language that
  consume the tool's stdout. The output format is small
  and documented in
  [`CAPABILITIES.md ## Observability`](CAPABILITIES.md#observability);
  if the user wants to script against it, the right
  answer is *"read the live source, write the parser
  against your installed binary"*.
- **A `samples/`, `bindings/`, or `reference/` subtree.**
  This is a thin loader for a shipped tool tree;
  substantive material lives in the source tree and in
  the GPUNetIO library docs.

## Loading order

1. Read this `SKILL.md` first to confirm the user's
   question is in scope (the user actually wants to
   measure sustained kernel-initiated WRITE BW through
   GPUNetIO, not learn GPUNetIO as a library or do a
   CPU-initiated measurement).
2. **For what the tool measures, the surface-selection
   rule against the GPI sister tool and the CPU-initiated
   `perftest`, the throughput-decomposition guide, the
   version overlay, the error taxonomy, the observability
   surface, and the safety overlay, see
   [CAPABILITIES.md](CAPABILITIES.md).**
3. **For step-by-step workflows — `install`, `configure`,
   `build`, `modify`, `run`, `test`, `debug`, `use` — see
   [TASKS.md](TASKS.md).**

## Related skills

- [`../../libs/doca-gpunetio/SKILL.md`](../../libs/doca-gpunetio/SKILL.md) —
  the library this tool wraps. The per-GPU `doca_gpu`
  context, the GPU-visible `doca_gpu_eth_*` and RDMA-side
  handles, the CUDA-side persistent-kernel pattern, the
  dual capability-discovery rule (DOCA cap-query AND
  `cudaGetDeviceProperties`), and the env preconditions
  (`nvidia_peermem` loaded, CUDA buffers registered with
  DOCA) live there.
- [`../../libs/doca-rdma/SKILL.md`](../../libs/doca-rdma/SKILL.md) —
  the underlying RDMA library. The RDMA queue this tool
  binds is created and connected via `doca-rdma`; the
  queue lifecycle, transport type (RC vs UC vs UD),
  permission matrix, and connection method are owned
  there.
- [`../../libs/doca-verbs/SKILL.md`](../../libs/doca-verbs/SKILL.md) —
  the raw-verbs escape hatch beneath `doca-rdma` /
  `doca-gpunetio`. This tool stays on the higher-level
  surfaces; `doca-verbs` is the right place only if the
  user needs a specific WR flag / QP attribute the
  GPUNetIO + RDMA surfaces do not expose.
- [`../doca-gpunetio-ib-write-lat/SKILL.md`](../doca-gpunetio-ib-write-lat/SKILL.md) —
  the latency analog of this tool. Same physical
  operation; same runtime framework; different metric
  class (BW vs latency). The two together carry the
  full GPUNetIO-side throughput / latency picture.
- [`doca-gpi`](../../libs/doca-gpi/SKILL.md) — the GPI
  programming surface (CUDA-kernel-initiated RDMA), the
  alternative runtime framework for the same physical
  operation. `doca/tools/` ships no GPI `ib_write_lat` /
  `ib_write_bw` benchmark binary, so the GPI comparison is
  against the library surface, not a sibling tool. The
  selection rule in
  [`CAPABILITIES.md ## Capabilities and modes`](CAPABILITIES.md#capabilities-and-modes)
  is the decision aid.
- [`doca-version`](../../doca-version/SKILL.md) — the
  canonical version-detection chain, four-way match rule,
  NGC container semantics, and headers-win-over-docs
  rule. The `## Version compatibility` section in this
  skill is a thin overlay; the body lives there.
- [`doca-setup`](../../doca-setup/SKILL.md) — env
  preparation, install verification, GPU + CUDA Toolkit
  pairing, `nvidia_peermem` load, hugepages, NUMA, and
  the *I have no install yet* path with the public NGC
  DOCA container.
- [`doca-public-knowledge-map`](../../doca-public-knowledge-map/SKILL.md) —
  routing to the public DOCA documentation set (DOCA GPU
  NetIO, DOCA RDMA pages on `docs.nvidia.com`) and the
  `docs.nvidia.com/cuda/` pointer for the CUDA Toolkit.
- [`doca-debug`](../../doca-debug/SKILL.md) — the
  cross-cutting debug ladder. The tool surfaces its own
  error taxonomy; when the cause is below DOCA, the
  taxonomy hands off here.
- [`doca-hardware-safety`](../../doca-hardware-safety/SKILL.md) —
  the bundle-wide hardware-safety meta-policy. The
  `## Safety policy` overlay cross-links it.