Network Troubleshooting
Layer-by-layer: ping for IP, traceroute for routing, dig for DNS, curl -v for HTTP — solves 90% of 'site is down' alerts.
Introduction
Layer-by-layer: ping for IP, traceroute for routing, dig for DNS, curl -v for HTTP — solves 90% of 'site is down' alerts.
Beginner analogy: think of Unix as a kitchen. The shell is the chef who reads your order, the kernel is the stove and fridge that actually cook and store, files are the ingredients, and pipes are the conveyor belts moving food from one chef to the next. Every Unix command you learn is one well-designed kitchen tool.
In this lesson we will walk through Network Troubleshooting step by step, see exactly how Linux handles it under the hood, look at the practical commands you will type every day on real servers, study a real-world DevOps scenario, and finish with the interview questions you will absolutely face when applying to AWS, Google Cloud, Red Hat, Netflix, Stripe and every modern infrastructure team.
Understanding the topic
Core concepts to understand:
- 🧠 Clear definition and mental model of network troubleshooting.
- 🐧 How the Linux kernel and shell collaborate to make it happen.
- 📂 Where files, processes and configuration live on a standard Linux server.
- 🔁 How network troubleshooting fits inside scripts, cron jobs and CI/CD pipelines.
- 🛡 Permissions, users, groups and least-privilege practices around network troubleshooting.
- 🚧 Common pitfalls: missing quotes, unset variables, wrong exit codes, dangerous
rm -rf. - 🏢 Real production scenarios at AWS, Netflix, Stripe and modern SaaS DevOps teams.
Syntax reference
Visual workflow / architecture:
User Command|vShell Interpreter|vCommand Parsing|vKernel Interaction|vSystem Resources|vCommand Execution|vTerminal Output
curl https://api.example.comApplication opens a socket via libc.
Informative example
Hands-on commands you can copy-paste:
Networking on Linux is a flat toolbox: ping tests reachability, curl -I hits an HTTP endpoint, ss (modern netstat) shows what is listening, and traceroute reveals the hops between you and the server.
# Quick network diagnosticsping -c 3 google.comcurl -I https://api.example.comss -tulpn | grep :80traceroute google.com | head -5
Sample terminal output:
64 bytes from 142.250.190.14: icmp_seq=1 ttl=117 time=12.3 msHTTP/2 200LISTEN 0 511 *:80 *:* users:(("nginx",pid=1023,fd=6))
Walk-through: notice how every Unix tool prints structured text and returns an exit code (0 = success, anything else = failure). That is the contract that lets you chain commands with &&, pipe them with |, and trust them inside automation. Reading these messages carefully is the difference between a senior Linux engineer and a junior one.
Real-world use
In production, Network Troubleshooting is part of every infrastructure engineer's daily flow at companies like AWS, Google Cloud, Netflix, Stripe, Shopify, GitHub and Red Hat. Engineers SSH into Linux servers, write small focused bash scripts, schedule them with cron or systemd timers, monitor them in Grafana and ship them through CI/CD. Mastering network troubleshooting means safer deploys, faster incident response and dramatically fewer 3 AM pages.
Best practices
- Always start scripts with
#!/bin/bashandset -euo pipefailso they fail fast on errors and unset variables. - Quote variables:
"$file"not$file— protects against spaces and word-splitting bugs. - Use absolute paths in cron, scripts and systemd units —
$PATHis minimal in those environments. - Log to
/var/log/<app>/and rotate withlogrotateso disks never fill up. - Run as the least-privileged user; reserve
sudofor the few commands that truly need root.
Common mistakes
rm -rf $VAR/when$VARis empty — wipes the whole filesystem. Always quote and validate.- Cron jobs that run from a fresh shell with no
$PATH— your script works manually but fails at 2 AM. - Forgetting
2>&1on logs — silent failures because stderr was thrown away. - Editing config files without taking a backup (
cp file file.bak) — no way to roll back.
Hands-on exercise
Interview preparation — practice these questions:
- Q1. Explain Network Troubleshooting in one sentence as if to a junior teammate.
- Q2. Walk through the exact Linux commands you would run for network troubleshooting on a production server.
- Q3. What is the difference between Unix and Linux, and where does network troubleshooting live in the stack?
- Q4. How would network troubleshooting behave inside a cron job vs an interactive shell, and why?
- Q5. Name two security or permission concerns around network troubleshooting and how you would mitigate them.
- Q6. How does network troubleshooting integrate with monitoring, logging and a CI/CD pipeline?
- Q7. Scenario: a 3 AM PagerDuty alert says network troubleshooting failed in production. Walk me through your debugging.