Systems Administration & Infrastructure

Systems administration is the craft of operating the computers and shared services an organization depends on. A system administrator turns raw machines and operating systems into maintained services with controlled access, predictable configuration, useful telemetry, and a tested recovery path.

The discipline spans on-premises equipment, virtual machines, cloud instances, employee endpoints, directory services, storage, and automation. The tools change; the operating principles do not.

TL;DR

Learning Tracks

Operating Systems

Learn processes, services, filesystems, permissions, logs, package management, and startup behavior across Linux, Windows, and macOS. Begin with Linux and the Command Line.

Identity & Access

Study directories, users, groups, roles, service accounts, federation, single sign-on, multifactor authentication, and privileged access. Continue with SSO, MFA, and Zero Trust.

Servers, Virtualization & Storage

Understand physical hosts, hypervisors, virtual machines, filesystems, network storage, capacity planning, snapshots, and the difference between availability and recoverability.

Configuration & Automation

Replace undocumented console work with versioned, reviewable automation. Ansible, Infrastructure as Code, and Terraform show how repeatability scales operations.

Maintenance & Recovery

Build patch rings, maintenance windows, monitoring, backup policies, restore drills, and disaster-recovery procedures. Use Monitoring, Alerting, and High Availability as companion topics.

The Administrator's Checklist

Start Here

  1. Build a small Linux virtual machine and document its users, services, ports, and update process.
  2. Configure key-based administrative access and remove unnecessary privileges.
  3. Automate the machine's baseline configuration.
  4. Add monitoring and centralized logs.
  5. In a disposable lab, verify the backup and rebuild a replacement VM from it. Validate restored data and services before removing the original lab VM.

Featured Topics

Operating Systems

Virtualization & Storage

Hardware

The server is replaceable; the directory state, privileged access, backups, and recovery design are not. Protect those deliberately.

Operating Principles

Keep systems boring, observable, reproducible, and recoverable. Define desired state in code where practical, separate administrative identities from everyday accounts, minimize exposed services, and make routine maintenance safe enough to perform regularly. Favor small, reversible changes with prechecks and post-change verification.

Troubleshooting Method

Establish the timeline and blast radius, check recent changes, compare healthy and unhealthy systems, then test one hypothesis at a time. Preserve logs and evidence before restarting services. After restoration, record the cause, detection gap, recovery path, and preventive action so the same failure becomes easier to diagnose.

Related Hubs

References