21 โ Linux: the ground everything runs on¶
Core question: Every container, pod, pipeline, and server ultimately runs on a Linux process โ can you navigate, interrogate, and fix one under pressure?
โฑ๏ธ Time: ~50 min ยท ๐๏ธ Level: BeginnerโIntermediate ยท ๐ Pehle chahiye: 00a Pre-flight
Is chapter ke baad tum kar paoge: - Navigate, read, and edit files and directories with confidence in any Linux environment - Debug a slow or unresponsive server using the golden five-command flow - Use pipes and text tools to mine any log file for errors, counts, and top offenders
How to use: ek terminal khol ke har command khud chalao โ reading โ knowing.
โฉ๏ธ Recall gate¶
Pehle memory se jawab do, phir kholo.
- (00a) What is a shell โ and what is the difference between a shell and a terminal?
- (00a) What is the difference between an absolute path and a relative path?
- (00a) What is
stdoutโ and how is it different fromstderr?
Jawab dekho
A terminal is the window you type in. A shell is the program running inside it (Bash, Zsh, sh) that interprets commands and returns output. The terminal is UI; the shell is the interpreter.
An absolute path starts from root
/and works from anywhere:/etc/nginx/nginx.conf. A relative path is relative to your current directory:./scripts/deploy.sh.pwdshows your current position.stdout (file descriptor 1) carries normal output. stderr (file descriptor 2) carries error messages. Both print to the screen by default. You can separate them:
cmd 2>/dev/nullsilences errors;cmd 2>&1merges both into stdout for capture.
Navigation & finding things¶
First thing you do on any unknown server: orient yourself.
| Command | What it does | Why you need it |
|---|---|---|
pwd |
Print current directory | Baseline โ where am I? |
ls -lah |
Long, all (hidden files), human-readable sizes | See everything including dotfiles |
cd /var/log |
Change directory | Move around the filesystem |
cd - |
Go to previous directory | Toggle between two locations |
cd ~ |
Go home | Back to your user's home directory |
tree -L 2 |
Visual tree, 2 levels deep | Quick structural overview of a project |
find / -name "*.log" 2>/dev/null |
Walk the tree, filter by name, suppress errors | Find any file when you don't know where it lives |
find . -type f -mtime -1 |
Files modified in the last 24 hours | What changed recently? |
find . -size +100M |
Files bigger than 100 MB | Disk-full culprit hunt |
which kubectl |
Full path of a binary | Is it installed? Which one am I running? |
locate nginx.conf |
Instant index-based search | Fast โ but the index (updatedb) may be stale |
๐ค Interview one-liner: "
findwalks the filesystem live with filters โ name, size, mtime, type.locateis instant but uses a prebuilt index that may lag. When the disk is full,find . -size +100Mis your first move."๐ฎ๐ณ Hinglish intuition:
findek real-time detective hai โ koi bhi file dhundh ke deta hai.locateek phone directory hai โ tez hai par puraani ho sakti hai. Disk full?find . -size +100Mseedha chor ko pakadta hai.
Viewing & searching content¶
Log files are how servers talk to you. Reading them well is half of incident response.
| Command | What it does |
|---|---|
cat file |
Dump whole file to screen (small files only) |
less file |
Paginated view โ /pattern to search, q to quit |
head -n 20 file |
First 20 lines |
tail -n 20 file |
Last 20 lines |
tail -f app.log |
Follow live โ new lines appear as they are written |
grep "ERROR" app.log |
Filter lines matching a pattern |
grep -ri "timeout" /etc |
Recursive, case-insensitive โ search a whole directory |
grep -c "500" access.log |
Count matching lines only |
grep -A3 -B1 "panic" app.log |
3 lines after + 1 before each match โ stack-trace context |
๐ค Interview one-liner: "
tail -fwatches a log live while a bug happens.grep -A3 -B1gives context around the match โ critical for stack traces where the useful line is two lines above the word ERROR."๐ฎ๐ณ Hinglish intuition:
tail -fek live TV channel hai.grepek remote hai โ sirf wahi dikhao jo chahiye. Dono saath mein:tail -f app.log | grep ERRORโ live error-only feed, baaki sab mute.
File operations¶
| Command | What it does | Notes |
|---|---|---|
mkdir -p a/b/c |
Create nested directories in one shot | -p creates parents if missing |
touch file.txt |
Create empty file or update timestamp | |
cp -r src/ dst/ |
Copy directory recursively | |
mv old new |
Move or rename โ works across paths | |
rm -rf dir/ |
Delete recursively, no confirmation prompt | No undo โ be certain before running |
ln -s /real/path link |
Create a symbolic link | Used for config files and versioned binaries |
โ ๏ธ
rm -rfhas no recycle bin. A misplacedrm -rf /orrm -rf ./*on a production server is irreversible. Alwayspwdandlsbefore running it.
Permissions โ know this cold¶
Every file on Linux has three permission sets โ owner, group, others โ each with three bits: read (4), write (2), execute (1).
-rwxr-xr--
โโโโโคโโโโคโโโโค
โโ โโ โโโโ others: r-- = 4
โโ โโโโโ group: r-x = 5
โโโโโ owner: rwx = 7
โโโ file type: - = regular, d = directory, l = symlink
The addition rule: 7 = 4+2+1 (rwx) ยท 6 = 4+2+0 (rw-) ยท 5 = 4+0+1 (r-x) ยท 4 = 4+0+0 (r--)
| Mode | Symbolic | Owner | Group | Others | Standard use |
|---|---|---|---|---|---|
755 |
rwxr-xr-x | rwx | r-x | r-x | Scripts, directories |
644 |
rw-r--r-- | rw- | r-- | r-- | Config files, static assets |
600 |
rw------- | rw- | --- | --- | SSH private keys, secrets |
777 |
rwxrwxrwx | rwx | rwx | rwx | Never in production |
chmod 755 script.sh # make a script executable by all users
chmod +x script.sh # symbolic: add execute bit to existing permissions
chmod -R 644 config/ # set 644 recursively on all files in a directory
chown devuser:appgroup file # change owner and group
chown -R devuser app/ # recursive ownership change
umask # show the default permission mask for new files
๐ค Interview one-liner: "r=4, w=2, x=1. 755 = owner full access, everyone else read+execute โ standard for scripts. 644 = owner read-write, everyone else read-only โ standard for config files. chmod 777 is a security hole: any process on the system can overwrite or execute the file."
๐ฎ๐ณ Hinglish intuition: Teen groups ke liye teen daraaze. r=4 padhna, w=2 likhna, x=1 chalana. 755 matlab owner ke paas master key hai; baaki sirf ander aa sakte hain โ kuch change nahi kar sakte.
Users, groups & sudo¶
| Command | What it does |
|---|---|
whoami |
Current username |
id |
UID, GID, and all group memberships |
sudo command |
Run one command as root |
su - user |
Switch to another user (full login environment) |
sudo useradd -m user |
Create a user with a home directory |
sudo passwd user |
Set or change a password |
sudo usermod -aG docker user |
Add user to a group (docker, sudo, etc.) |
groups user |
List all groups a user belongs to |
cat /etc/passwd |
All user accounts on the system |
cat /etc/group |
All groups and their members |
๐ค Interview one-liner: "
sudo usermod -aG docker useradds a user to the docker group so they can run Docker withoutsudo. The-aflag appends โ omitting it would replace all group memberships, locking the user out of other groups."
Processes¶
A process is any running program. Every container, every service, every script is a process with a PID (Process ID). Knowing how to find and control processes is the core of incident response.
| Command | What it does |
|---|---|
ps aux |
Snapshot of all processes โ user, PID, CPU%, MEM%, command |
ps aux \| grep nginx |
Find a specific process by name |
top |
Live monitor โ CPU and memory per process, updated every second |
htop |
Colour-coded top with mouse support (install separately) |
kill <PID> |
Send SIGTERM (15) โ polite shutdown request |
kill -9 <PID> |
Send SIGKILL โ forced, immediate, no cleanup |
pkill -f "python app" |
Kill all processes matching a name or pattern |
nohup ./run.sh & |
Run in background, survives terminal close |
jobs / fg / bg |
List / foreground / background shell jobs |
Signal numbers every engineer must know:
| Signal | Number | Meaning |
|---|---|---|
| SIGHUP | 1 | Reload config โ most daemons re-read their config on kill -1 without restarting (nginx keeps the master PID, spawns new workers, drains old ones) |
| SIGTERM | 15 | Polite shutdown โ app can flush data, close connections, then exit |
| SIGKILL | 9 | Forced kill by the kernel โ no cleanup, no ceremony |
๐ค Interview one-liner: "Always try SIGTERM first โ just
kill <PID>. SIGKILL is the last resort because it skips cleanup: open files can be left corrupt and database connections abandoned."Containers are Linux processes wrapped in namespaces and cgroups โ see M3 Docker for how Docker uses these primitives, and M9 for how Kubernetes orchestrates them at scale.
System & resources โ the debug flow¶
This is the section most directly tied to incident response. Each command answers one specific diagnostic question.
| Command | What it shows | Incident question it answers |
|---|---|---|
top |
CPU and memory per process, load average, system totals | "What process is burning CPU or RAM?" |
df -h |
Disk used/free per filesystem, human-readable | "Is the disk full?" |
df -i |
Inode usage per filesystem | "Disk shows space but writes fail?" |
du -sh * |
Size of each item in the current directory | "Which directory is the bloat?" |
du -sh /var/log |
Total size of a specific path | "How big are my logs?" |
free -h |
RAM: total, used, free, buffers/cache, swap used | "Are we swapping? Is OOM near?" |
uptime |
Load average (1 / 5 / 15 min) + uptime | "Is load > core count?" (overloaded) |
nproc |
Number of CPU cores | Context for interpreting load average |
uname -a |
Kernel version, hostname, architecture | "What OS and kernel is this exactly?" |
๐ค Interview one-liner: "When a server is slow or down I run
top(CPU/mem spike?),df -h(disk full?),free -h(swapping?),journalctl -xe(recent errors), andss -tulnp(port conflict?). That five-command sequence covers 90% of real incidents."โ ๏ธ The inode trap:
df -hcan show gigabytes free while the disk is "full" โ if you have millions of tiny files (log fragments, temp files, cache entries), inodes (filesystem metadata slots) can be exhausted even with free block space. Always checkdf -iwhendf -hlooks healthy but writes are failing.
See M8 Observability for the monitoring-layer view of the same problems โ what you are doing manually here, Prometheus + Grafana does automatically and continuously.
Networking¶
Every connection, every port, every DNS query is inspectable from the command line.
| Command | What it does | When to use it |
|---|---|---|
ip a |
All network interfaces and assigned IPs | "What IP does this host have?" |
ss -tulnp |
All listening sockets: protocol, port, PID, process name | "What is listening on port X?" |
ping host |
ICMP reachability โ round-trip time | "Is this host alive on the network?" |
curl -I https://site |
HTTP headers only โ shows status code fast | "Is the web server responding?" |
curl -v http://svc:8080 |
Full verbose request + response trace | "Why can't I reach this service?" |
wget url |
Download a file over HTTP/HTTPS | Fetch artefacts, test download paths |
dig example.com |
Full DNS resolution chain | "Is DNS resolving correctly?" |
nslookup example.com |
Simple DNS query | Quick lookup โ faster to type than dig |
nc -zv host 5432 |
Test TCP port reachability without sending data | "Can I reach the database port?" |
traceroute host |
Each routing hop on the path | "Where is the packet being dropped?" |
๐ค Interview one-liner: "
ss -tulnpis the first command for 'port already in use' โ it shows every listening port and the PID that owns it.curl -vandnc -zvtest whether a remote port is actually reachable, which separates network/firewall problems from application problems."๐ฎ๐ณ Hinglish intuition:
ss -tulnpek building directory hai โ kaun sa process kaun si window pe baith ke sun raha hai.nc -zv host portek door knock hai โ khuli hai ya nahi.curl -vek poori conversation hai โ andar jao, poochho, jawab suno.
Packages & services¶
| Command | What it does |
|---|---|
sudo apt update |
Refresh package index (Debian/Ubuntu) |
sudo apt install -y nginx |
Install without interactive confirmation |
sudo apt remove nginx |
Uninstall a package |
apt list --installed |
List all installed packages |
sudo dnf install -y nginx |
Install on RHEL / Fedora / Amazon Linux |
sudo systemctl start nginx |
Start a service right now |
sudo systemctl stop nginx |
Stop a service |
sudo systemctl restart nginx |
Stop + start (pick up config changes) |
sudo systemctl enable nginx |
Start automatically on every boot |
sudo systemctl status nginx |
Current state + last few log lines |
journalctl -u nginx -f |
Follow a service's log output live |
journalctl -u nginx --since "10 min ago" |
Recent logs for a specific service |
journalctl -xe |
System-wide recent errors with context |
๐ค Interview one-liner: "
systemctl enablemeans 'start on every boot'.systemctl startmeans 'start right now'. You need both on a new server.journalctl -u <service> -fis the first look when a service won't start โ it shows the exact error before the process died."For the error-response reflex for common service failures see 16 Appendix. For the full observability stack built on top of these logs see M8.
Text processing & pipes โ the Unix superpower¶
Unix philosophy: small tools, each doing one thing, chained with |. The combination is more powerful than any single application.
# Count 500 errors in an access log
cat access.log | grep "500" | wc -l
# Frequency table โ most common items first
sort file | uniq -c | sort -rn
# Top IPs hitting a web server (field 1 of Apache/nginx access log)
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10
# Find-and-replace across a file
sed 's/old_value/new_value/g' config.template > config.final
# Extract a column by delimiter (field 1 of /etc/passwd = usernames)
cut -d: -f1 /etc/passwd
# Pull a specific field from structured log lines
grep "ERROR" app.log | awk '{print $4}'
| Tool | One job |
|---|---|
grep |
Filter lines by pattern |
awk |
Extract columns / arithmetic on fields |
sed |
Find-and-replace in a byte stream |
cut |
Extract columns by delimiter |
sort |
Sort lines lexicographically or numerically |
uniq -c |
Count consecutive duplicates (always sort first) |
wc -l |
Count lines |
head -N / tail -N |
First / last N lines of a stream |
๐ค Interview one-liner: "
sort | uniq -c | sort -rn | headis the frequency-table recipe โ it turns any log into a ranked list of top errors, IPs, or status codes in one pipeline. Stick it on any stream."๐ฎ๐ณ Hinglish intuition: Pipes ek assembly line hain.
grepfilter karta hai,awkcolumn nikalti hai,sort|uniq -cginne wala hai,sort -rnbada pehle laata hai. Chain bano โ bada kaam chhota ho jaata hai.
Archives & transfer¶
# Create a compressed archive
tar -czf backup.tar.gz dir/
# Extract an archive
tar -xzf backup.tar.gz
# List archive contents without extracting
tar -tzf backup.tar.gz
# Copy a file to a remote host over SSH
scp file.txt user@host:/remote/path/
# Sync a directory (only changed files, preserves permissions)
rsync -avz src/ user@host:/dst/
tar flags decoded: -c create ยท -x extract ยท -z gzip compression ยท -f filename follows ยท -v verbose output.
rsync -avz is preferred over scp for directories โ it sends only diffs, preserves attributes (-a), compresses in transit (-z), and shows progress (-v). On large deployments the bandwidth savings are significant.
Environment & shell¶
echo $PATH # colon-separated directories the shell searches for commands
export API_KEY=abc # set an env var visible to this shell and all child processes
env # print every environment variable currently set
alias ll='ls -lah' # create a shorthand (add to ~/.bashrc to persist across sessions)
history # numbered list of past commands
!123 # re-run history entry 123
Ctrl+R # reverse-search history โ type part of a past command
man ls # full manual page (q to quit)
ls --help # short inline help for most commands
โ ๏ธ
export VAR=valuesets the variable for the current session only. To persist, add it to~/.bashrc(user) or/etc/environment(system-wide). Withoutexport, child processes (scripts, subshells) cannot see the variable โ a common source of "env var works in terminal but not in my script" confusion.
The golden debug flow¶
When someone says "the server is slow" or "my service is down", this sequence covers 90% of real incidents. Run them in order โ each either finds the problem or rules it out.
# Step 1 โ Is CPU spiking? Which process owns it?
top
# Press 'P' to sort by CPU usage, 'M' for memory, 'q' to quit.
# Watch for a process at 99% CPU or total load > nproc output.
# Step 2 โ Is the disk full?
df -h
# Any filesystem at 100% Use%? Disk full = writes fail = services crash.
# Drill down: du -sh /var/* 2>/dev/null | sort -rh | head -10
# Step 2b โ Are inodes exhausted? (disk "full" with space remaining)
df -i
# If IUse% is 100% on any filesystem, you have inode exhaustion โ not block exhaustion.
# Find the flood: find /tmp -type f | wc -l
# Step 3 โ Are we out of memory or swapping heavily?
free -h
# Check the 'available' column. Near zero + swap in use = OOM is imminent.
# Kubernetes calls this OOMKilled (exit code 137).
# Step 4 โ What did the service log right before it died?
journalctl -xe
# Or for a specific service:
journalctl -u myservice --since "15 min ago"
# Step 5 โ Is something unexpected listening, or is the port already taken?
ss -tulnp
# Shows every listening port and the PID + process name that owns it.
The decision tree:
top
โโโ CPU/mem spike found? โ identify PID, kill or investigate
โโโ No spike?
df -h
โโโ Filesystem at 100%? โ du -sh to find bloat, delete or rotate
โโโ Looks ok?
df -i
โโโ IUse% at 100%? โ find and delete the tiny-file storm
โโโ Looks ok?
free -h
โโโ Available near zero? โ OOM risk; reduce load or add memory
โโโ Ok?
journalctl -xe
โโโ Error found? โ fix the root cause
โโโ No obvious error?
ss -tulnp
โโโ Wrong port or unexpected listener? โ kill/reconfigure
โโโ All clear โ application-level bug; check app logs
๐ฎ๐ณ Hinglish intuition: Yeh flow ek doctor ka checkup hai. Pehle pulse (
top), phir breathing (df -h), phir blood pressure (free -h), phir patient ki history (journalctl), phir sunna stethoscope se (ss -tulnp). Sab normal? Toh application ka bug hai โ logs mein jao.
This debug flow is the manual version of what M8 Observability automates at scale. See the connected system diagram for how the Linux layer relates to every layer above it.
Small but important¶
Commonly missed in tutorials, constantly used in real work:
| Pattern | What it does | Why it matters |
|---|---|---|
cmd1 && cmd2 |
Run cmd2 only if cmd1 exits 0 (success) |
apt update && apt install โ don't install from a stale index |
cmd1 \|\| cmd2 |
Run cmd2 only if cmd1 fails |
Fallback / default logic in scripts |
cmd1 ; cmd2 |
Run cmd2 always, regardless of cmd1 outcome |
Fire-and-forget sequencing |
echo hi > f |
Overwrite file with stdout | Creates or truncates โ destructive |
echo hi >> f |
Append stdout to existing file | Add without destroying prior content |
cmd 2>&1 \| tee out.log |
Merge stderr into stdout, print live AND save to file | Capture full output during long runs |
find . -name "*.log" \| xargs rm |
Feed output lines as arguments to another command | Bulk operations when a glob won't reach |
watch -n2 'df -h' |
Re-run a command every 2 seconds | Monitor disk / memory without polling manually |
df -i |
Show inode usage | A disk can be "full" on inodes with free block space |
crontab -e |
Edit the current user's cron schedule | Recurring jobs: MIN HOUR DOM MON DOW command |
sudo !! |
Re-run the previous command with sudo prepended | "Permission denied" โ sudo !! saves retyping the whole command |
Signal quick-reference:
| Signal | Number | kill syntax |
App reaction |
|---|---|---|---|
| SIGHUP | 1 | kill -1 <PID> |
Reload config (most daemons) |
| SIGTERM | 15 | kill <PID> |
Graceful shutdown โ flush, close, exit cleanly |
| SIGKILL | 9 | kill -9 <PID> |
Immediate forced kill โ no cleanup possible |
Hands-on lab¶
โ
Prove it โ bash labs/check-m0-linux.sh
Lab ho gaya? Tick mat lagao โ machine se verify karo (redirection ยท chmod +x ยท chmod 644 ยท awk field extract ยท ss port snapshot). โ pe exact fix-hint milega. โ The Doer's Path
No cloud account needed. Use WSL, a local VM, or docker run -it ubuntu bash.
Drill 1 โ Disk-full investigation
# Where is disk space going?
df -h
# Find the biggest directories under /var
du -sh /var/* 2>/dev/null | sort -rh | head -10
# Find any file over 100 MB
find /var/log -size +100M 2>/dev/null
# Check inodes too
df -i
Drill 2 โ Service down diagnosis
# Check a service's current state
sudo systemctl status nginx
# Read its recent log output
journalctl -u nginx --since "10 min ago"
# If nginx is not installed, list running services instead
systemctl list-units --type=service --state=running | head -10
Drill 3 โ Port conflict
# What is listening on port 8080?
ss -tulnp | grep :8080
# What is listening on all ports?
ss -tulnp
# Cross-reference a PID to a process name
ps aux | grep <PID>
Drill 4 โ CPU and memory snapshot
# Open top โ press 'P' (CPU sort), 'M' (memory sort), 'q' (quit)
top
# Memory breakdown
free -h
# Load average vs core count โ if load-1min > nproc, system is overloaded
uptime
nproc
Drill 5 โ Log mining with pipes
# Count 404 errors in nginx access log
grep "404" /var/log/nginx/access.log 2>/dev/null | wc -l
# Frequency table of HTTP status codes (field 9 in nginx combined format)
awk '{print $9}' /var/log/nginx/access.log 2>/dev/null \
| sort | uniq -c | sort -rn
# If no web server, scan auth log for failed logins
grep "Failed" /var/log/auth.log 2>/dev/null | tail -20
โ Sahi hua to aisa dikhega:
- Drill 1:
df -hprints a table withUse%per filesystem.find /var/log -size +100Meither returns file paths (you found the culprits) or prints nothing (no large files โ both are valid outcomes).df -ishows inode usage โ anything near 100% is a problem worth investigating. - Drill 3:
ss -tulnp | grep :8080returns a line including a PID and process name if something is listening on 8080, or no output if the port is free. You can identify the process from the PID column. - Drill 4:
topopens a live display; sorting by CPU/memory puts the heaviest process at the top.uptimeshows three load numbers โ compare the first tonproc. Load exceeding core count means the system cannot service all queued work. - Drill 5: Each
awk | sort | uniq -c | sort -rnpipeline produces a ranked frequency table. If no matching log file exists, the command exits cleanly with no output โ that is also a correct result.
Summary¶
| Topic | The one thing to remember |
|---|---|
| Navigation | find . -size +100M for disk culprits; ls -lah to see everything including hidden files |
| Permissions | r=4, w=2, x=1 ยท 755 for scripts ยท 644 for configs ยท 600 for secrets ยท never 777 |
| Processes | kill = SIGTERM (graceful) ยท kill -9 = SIGKILL (last resort) ยท ps aux \| grep to find PIDs |
| Debug flow | top โ df -h โ df -i โ free -h โ journalctl -xe โ ss -tulnp covers 90% of incidents |
| Networking | ss -tulnp = who is listening where ยท curl -v + nc -zv = is the port reachable |
| Services | enable = on boot ยท start = right now ยท journalctl -u svc -f = live logs |
| Pipes | sort \| uniq -c \| sort -rn \| head = frequency table from any byte stream |
| Gotchas | df -i for inodes ยท && vs ; in scripts ยท export makes vars visible to children ยท sudo !! to retry |
Self-check quiz¶
Pehle memory se jawab do.
- What does
chmod 644 file.txtallow โ and who can write to that file? - You run
kill 1234but the process does not die. What do you try next, and why is it a last resort? df -hshows 2 GB free on/var. Your application still cannot write a file there. What do you check?- A service crashed 3 minutes ago. What single command shows its log output from the last 5 minutes?
ss -tulnpshows a process on port 8080 that should not be there. What are your next two steps?- You need a command to run every night at 2 AM. What tool do you use, and what does the schedule field look like?
- What is the difference between
cmd1 && cmd2andcmd1 ; cmd2?
Jawab dekho
-
644= owner rw-, group r--, others r--. Only the owner can write. Group and others can only read. This is the standard for config files โ world-readable but not world-writable. -
Try
kill -9 <PID>(SIGKILL). It is a last resort because it bypasses the application's cleanup code โ open file handles may not be flushed, database connections not closed, in-flight transactions not rolled back. Use it only when the process genuinely refuses to terminate gracefully. -
Check inode exhaustion:
df -i. IfIUse%is at 100% on/var, the filesystem has no free inode slots even though block space is available. This happens when millions of small files accumulate. Find and delete the file storm:find /var -type f | wc -lto quantify, then investigate subdirectories. -
journalctl -u <service-name> --since "5 min ago"โ orjournalctl -u <service-name> -n 100for the last 100 lines regardless of time. -
First: identify โ
ps aux | grep <PID>to see what the process actually is. Second: act โ if it is a known conflicting service,systemctl stop <service>. If it is an unknown process, investigate before killing it. -
Use
crontab -e. An entry for 2 AM daily:0 2 * * * /path/to/script.sh. Fields left to right: minute (0), hour (2), day-of-month (), month (), day-of-week (*). -
&&is conditional โcmd2runs only ifcmd1exits with code 0 (success).;is unconditional โcmd2runs regardless. In scripts,&&prevents cascading failures;;is "do both no matter what." In CI pipelines&&is almost always what you want.
Interview questions¶
1. What is the difference between a soft link (symlink) and a hard link? A symlink points to a path โ if the original file moves or is deleted, the symlink breaks and returns "No such file or directory." A hard link points directly to the inode (the actual data on disk) โ the data persists until every hard link to it is removed. Hard links cannot span filesystems; symlinks can. In practice, symlinks are far more common and are used heavily for versioned binaries and config files.
2. Why is chmod 777 a security problem?
It grants every user on the system โ including any compromised web process, container escape, or rogue script โ full read, write, and execute access. Least-privilege principle: scripts should be 755 (executable, not writable by others), config files 644, secret files 600. 777 is the "I'll fix permissions later" setting that never gets fixed.
3. What is the difference between SIGTERM and SIGKILL?
SIGTERM (signal 15, default kill) is a polite request โ the process catches it, runs cleanup code, closes connections, flushes buffers, and exits gracefully. SIGKILL (signal 9, kill -9) is enforced directly by the kernel โ the process has no chance to clean up. Always try SIGTERM first. Use SIGKILL only when the process is truly hung.
4. How do you find the largest files on a filesystem?
Start with df -h to identify which filesystem is full. Then du -sh /path/* 2>/dev/null | sort -rh | head -10 to drill into directories. For individual files: find / -size +100M 2>/dev/null. The 2>/dev/null suppresses permission-denied noise.
5. How do you find what is listening on a specific port?
ss -tulnp | grep :8080. The output includes protocol, local address, PID, and process name. You can then ps aux | grep <PID> for more detail or kill <PID> to stop it.
6. What is the difference between an environment variable and a shell variable?
A shell variable (VAR=value) is local to the current shell โ child processes and scripts cannot see it. An environment variable (export VAR=value) is inherited by all child processes (scripts, subshells, programs launched from the shell). In Docker and Kubernetes, all runtime config is passed as environment variables โ which is why export matters so much in automation.
7. Where are system logs and what is the modern way to read them?
Traditional text logs live in /var/log/ (syslog, auth.log, kern.log). On all modern systemd-based Linux (Ubuntu 16+, RHEL 7+, Debian 8+), the primary store is the journal โ read with journalctl. Use journalctl -u <service> for a specific service, journalctl -xe for recent system errors, journalctl -f to follow live.
8. What does load average mean and when is it "too high"?
Load average (shown by uptime) is the average number of processes waiting for CPU time over the last 1, 5, and 15 minutes. Compare to nproc (core count). Load equal to core count = 100% busy but keeping up. Load greater than core count = processes are queueing โ the system cannot service demand. A 1-minute load of 8.0 on a 4-core machine means processes are waiting on average twice as long as they should.
20-second cheat-sheet¶
NAVIGATE pwd ls -lah cd /path cd - find . -size +100M which cmd
VIEW tail -f less grep -ri grep -A3 -B1 "pattern" file
PERMISSIONS r=4 w=2 x=1 | 755=scripts 644=files 600=secrets never 777
PROCESSES ps aux|grep top(P=cpu M=mem) kill(TERM) kill -9(KILL)
DEBUG FLOW top โ df -h โ df -i โ free -h โ journalctl -xe โ ss -tulnp
NETWORK ss -tulnp curl -v nc -zv host port dig traceroute
SERVICES systemctl start|enable|status journalctl -u svc -f
PIPES grep | awk | sort | uniq -c | sort -rn | head wc -l sed cut
ARCHIVES tar -czf out.tgz dir/ tar -xzf file.tgz rsync -avz src/ dst/
SMALL BITS && vs ; 2>&1|tee xargs watch -n2 df -i sudo !!