PyTorch Tensors & Autograd

Everything in PyTorch is built on two ideas. Tensors are n-dimensional arrays like NumPy's, but able to live on GPUs and other accelerators. Autograd is the engine that records operations on tensors and automatically computes gradients by backpropagation. Neural network layers, optimizers, and loss functions are all conveniences on top of these two primitives.

Understanding tensors (shapes, dtypes, devices, broadcasting, and views) prevents most day-to-day PyTorch bugs, like shape mismatches, silent broadcasting errors, and "expected all tensors to be on the same device" exceptions. Understanding autograd explains training: why you call zero_grad(), what .backward() does, and when to use no_grad().

TL;DR

Quick Example

Fitting a line with raw tensors and autograd, no nn modules:

The training loop page shows the idiomatic version with nn.Module and optimizers.

Core Concepts

Creating Tensors

Shapes, dtypes, and Devices

Broadcasting

When shapes differ, PyTorch aligns them from the trailing dimension and expands size-1 dimensions:

Use keepdim=True in reductions and assert shapes in critical code.

Views vs Copies

view, slicing, transpose, permute, and expand return views that share storage: modifying one modifies the other. clone() makes an independent copy. Some operations need contiguous memory: view fails after transpose, while reshape (copying if needed) or .contiguous().view(...) works. In-place operations (methods ending in _, like add_) save memory but can break autograd if they overwrite values needed for gradients.

Autograd

When a tensor has requires_grad=True (all nn.Parameters do), each operation records a node in a dynamic computation graph (define-by-run). Calling .backward() on a scalar (usually the loss) traverses the graph in reverse, applying the chain rule, and accumulates gradients into each leaf tensor's .grad.

Disabling Gradient Tracking

Best Practices

Be Explicit About Devices

Define a single device variable, move the model once, and create or move data per batch. Avoid .cuda() sprinkled through code, which breaks on CPU and Apple Silicon (MPS) machines.

Assert Shapes at Boundaries

Shape bugs often don't raise errors; broadcasting silently produces wrong results. Add assertions or use named dimensions in comments (# (B, T, C)) around model inputs, losses, and reshapes.

Use item() and detach() for Logging

Logging loss itself keeps the whole graph alive and leaks memory across iterations. Log loss.item() (a Python float) or loss.detach().

Prefer Out-of-Place Operations Unless Memory-Bound

In-place operations can cause autograd errors ("a leaf Variable that requires grad is being used in an in-place operation") or subtle bugs. Use them deliberately, for memory savings.

Common Mistakes

Forgetting to Zero Gradients

Call optimizer.zero_grad() (with set_to_none=True, the default in recent versions) each step, unless you're intentionally accumulating gradients.

Mixing Devices

RuntimeError: Expected all tensors to be on the same device usually means the model is on GPU but a batch, mask, or freshly created tensor is on CPU. Create auxiliary tensors with device=x.device.

Converting to NumPy Inside the Graph

Calling .numpy() on a tensor that requires grad raises an error, and round-tripping through NumPy mid-computation breaks gradient flow. Stay in torch operations for anything that needs gradients.

FAQ

What's the difference between a PyTorch tensor and a NumPy array?

Both are n-dimensional arrays with similar APIs. PyTorch tensors can run on GPUs and other accelerators, and they support automatic differentiation. CPU tensors and NumPy arrays can share memory via torch.from_numpy and .numpy().

What does loss.backward() actually do?

It walks the computation graph recorded during the forward pass in reverse, computing gradients of the loss with respect to every tensor with requires_grad=True, and adds them to those tensors' .grad attributes. The optimizer then uses these gradients to update parameters.

When should I use torch.no_grad() vs inference_mode()?

Both disable gradient tracking. inference_mode is stricter and slightly faster, so it's ideal for serving and evaluation where tensors never re-enter autograd. no_grad is more flexible, for example for manual parameter updates or when outputs might later be used with autograd.

Why do gradients accumulate instead of being overwritten?

Accumulation enables techniques like gradient accumulation across micro-batches (simulating bigger batches) and summing gradients from multiple losses. The cost is that you must explicitly reset them each optimization step.

Related Topics

References