Linux is everywhere, and it is one of the most important pieces of software humans have ever developed. It is an operating system that powers the majority of servers, embedded systems, and mobile devices.
These are my Linux notes as a backend developer. They cover what happens when the machine boots, how files, users, processes and services work, and the commands I reach for every day on servers and in containers. It is written for developers who use Linux daily but never sat down to learn it properly.
Terminology
Let’s start with some basic terminology:
- Kernel: The kernel acts as the core of the Linux operating system, managing the communication between hardware and software. It oversees hardware operations and enables applications to utilize these resources effectively.
- Process: A process is an instance of a program that is being executed. A process can be a daemon or a service, but it can also be a user-initiated program.
- Service: A service is a program that operates as a background process within the system.
- Daemon: A daemon is a type of service that runs in the background and is not directly controlled by the user. Daemons are usually started at boot time and are controlled by the init process, which is systemd on most distributions.
- Shell: The shell is the command line interpreter that provides a user interface for the user to interact with the kernel. It is a program that takes commands from the keyboard and gives them to the operating system to perform.
- The Filesystem: A filesystem is a method for storing and organizing files and directories on a storage device.
- Display server: The display server draws windows on the screen and passes input to applications. Current desktops use Wayland; the older X Window System (X11) is still around on older systems.
- Distribution: A Linux distribution is a collection of software that is built on top of the Linux kernel and provides additional tools and features.
Boot
When you switch on the computer, the firmware, UEFI on current machines and the Basic Input/Output System (BIOS) on older ones, begins initializing the hardware and conducts a test of the main memory. This routine is known as POST, which stands for Power On Self Test.
Once the POST is completed, system control passes from the firmware to the Bootloader (usually GRUB or systemd-boot), which is a small program that is
responsible for loading the operating system. On UEFI machines it lives on the EFI system partition, usually mounted at /boot/efi or /efi.
It starts by loading the Kernel Image and the initial RAM filesystem (initramfs) into memory (which contains some critical files and device drivers needed to start the system) and then transfers the control to the kernel.
After the kernel has configured all the hardware and mounted the root filesystem, it executes /sbin/init, which on almost every distribution today is a link to systemd. This process
then becomes the initial system process (PID 1), subsequently initiating other processes to operationalize the system.
While most processes on the system ultimately derive from init, exceptions exist in the form of kernel processes.
These are initiated directly by the kernel and are tasked with handling the internal aspects of the operating system.
Graphical session
Loading the graphical user interface is one of the final steps in the boot process. It allows users to interact with the system using graphical elements such as windows, icons, and menus.
A component known as the Display Manager shows the login screen, oversees graphical user logins and initiates the relevant desktop environment once a user has logged in. The default display manager for GNOME is gdm, and for KDE Plasma it is sddm.
On current distributions the desktop runs on Wayland, where the compositor (GNOME Shell or KWin) acts as the display server. Older applications written for X11 still work through XWayland.
A desktop environment comprises a session manager, which initiates and oversees the elements of the graphical session, and the window manager (on Wayland, the compositor), which manages the positioning and movement of windows, including their title-bars and other controls.
Virtual Terminals
They are called “virtual” because, although there can be multiple active terminals, only one terminal remains visible at a time.
One or two virtual terminals (usually VT 1 and 2, or VT 7 on older systems) are reserved for the graphical environment, and text logins are enabled on the unused VTs.
To switch between VTs, press CTRL-ALT-function key for the VT.
# for Virtual Terminal 6 press:
CTRL + ALT + F6Reboot and Shutdown
# power off in one minute, sends a warning message and prevents users from logging in
sudo shutdown
# power off now
sudo shutdown now
# power off at 10:00 with a message
sudo shutdown -h 10:00 "Shutting down for scheduled maintenance."
# cancel a scheduled shutdown
sudo shutdown -c
# reboot now
sudo shutdown -r now
sudo rebootFiles and the filesystem
Everything is a file!
Linux Directory Structure
/The root directory: All other directories are subdirectories of root, which is the top of the directory tree structure./binEssential user command binaries that need to be available for all users cat, ls, cp etc./bootStatic files of the bootloader./devDevice files such as terminal devices, usb or any device attached to the system./etcHost-specific system configuration./homeUser home directories./libEssential shared libraries and kernel modules./mediaMount point for removable media./mntMount point for mounting a filesystem temporarily./optAdd-on application software packages./sbinEssential system binaries./srvData for services provided by this system./tmpTemporary files./usrSecondary hierarchy for read-only user data; contains the majority of (multi-)user utilities and applications./varVariable data; files whose content is expected to continually change during normal operation of the system./rootHome directory for the root user./procVirtual filesystem providing process and kernel information as files.
On current distributions /bin, /sbin and /lib are symbolic links to /usr/bin, /usr/sbin and /usr/lib (the “merged /usr”), so the split only matters on older systems.
Moving around
# find current path
$ pwd
# change to user home directory
$ cd
$ cd ~
# go to parent directory
$ cd ..
# go back to the previous directory
$ cd -
# changes your current directory to the root (/) directory.
$ cd /
# list the contents of the present working directory.
$ ls
# list all files, including hidden files and directories.
$ ls -a
# list files with permissions, owner, size and date.
$ ls -l
# displays a tree view of the filesystem (install the tree package first).
$ tree
# add a directory to the top of the directory stack, changing the current working directory to that directory.
$ pushd /etc
# remove the directory at the top of the stack, changing the current working directory to the new top.
$ popd
# display the directory stack.
$ dirs
# modify timestamps of files or create new empty files.
$ touch file1.txtLinks
Links are essentially shortcuts or references to a file or directory.
Hard links
Hard links are simply additional names for an existing file.
$ ln /path/to/original /path/to/linkSoft links (Symbolic Links)
A symbolic link is like a shortcut to the original file or directory. If you delete the original file, the symbolic link will still exist but will point to a non-existing location.
$ ln -s /path/to/original /path/to/linkSearch for files
locate searches a database of file names, which is fast but only as fresh as the last updatedb run. On most distributions it comes from the plocate package.
# executes a search utilizing an existing database of system files and directories.
$ locate file1.txt
# get a shorter list of results using grep to print only the lines that contain one or more specified strings
$ locate file1 | grep test
# refresh the locate database
$ sudo updatedb
# search for files
$ find / -name file1.txt
# find files whose status changed 3 days ago
$ find / -ctime 3
# find files based on size
$ find / -size +100MDisk space
# free space on each mounted filesystem
df -h
# size of each file and directory in the current directory
du -sh *
# list disks and partitions
lsblkFilesystems and mounting
Linux filesystems are crucial for organizing and accessing data. Linux supports multiple filesystem types like ext4, XFS, Btrfs, FAT32, and NTFS, each with unique features.
The filesystems manage not only files but also metadata like permissions and support symbolic and hard links. Linux’s filesystem hierarchy integrates all storage into a unified directory tree, starting from the root directory (/).
Additionally, Linux supports network-based filesystems, notably NFS (Network File System), which allows files to be shared across a network. NFS enables multiple clients to access data stored on a server as if it were locally mounted, enabling seamless file sharing across different systems and networks.
Mounting is essential in Linux, enabling access to filesystems at specific points in this directory tree. Mounting refers to the process of making a filesystem accessible at a certain point in the directory structure. In short, you need to mount a filesystem to access its contents.
# mount the filesystem located on the partition /dev/sda5 to the directory /home
sudo mount /dev/sda5 /home
# unmount
sudo umount /homeIf you want it to be automatically available every time the system starts up, you need to edit /etc/fstab accordingly.
The proc filesystem
Some filesystems, such as those mounted at /proc, are referred to as pseudo-filesystems because they don’t have a physical representation on the disk.
The /proc filesystem is very important as it collects information on-demand without requiring disk storage.
You can find process statuses and other useful information.
# process 737 info
cat /proc/737/status
# cpu info
cat /proc/cpuinfoComparing Files with diff
# compare two files
diff file1.txt file2.txt
# a more readable format
diff -u file1.txt file2.txt
# side by side diff
diff -y file1.txt file2.txt
# recursive directory comparison
diff -r dir1 dir2
# add the differences into a patch file
diff -u file1.txt file2.txt > changes.patch
# apply the patches
patch file1.txt changes.patchArchives with tar
# create a gzip-compressed archive of a directory
tar -czf archive.tar.gz directory/
# list the contents of an archive
tar -tzf archive.tar.gz
# extract an archive
tar -xzf archive.tar.gz
# extract into another directory
tar -xzf archive.tar.gz -C /path/to/destinationRsync
Rsync stands for remote sync. It’s a utility tool for efficient file transfer and synchronization.
# rsync a file from remote host to local in archive mode (keeps permissions and timestamps), verbose and compressed
rsync -avz user@remote_host:/path/to/remote/file /path/to/local/destinationDisk To Disk Copy
dd overwrites the target completely, so check of= twice before you press enter.
# if specifies the input and of the output
# bs=64K sets the block size
# status=progress outputs the progress
sudo dd if=/dev/sda of=/dev/sdb bs=64K conv=noerror,sync status=progressUsers and permissions
Every file has an owner, a group and a set of permissions for the user (owner), the group and others.
ls -l shows them, for example -rwxr-xr-x 1 achilles achilles 1024 Oct 1 10:00 script.sh.
Users, groups and sudo
# who am I, and which groups am I in
whoami
id
# run a command as root
sudo command
# add yourself to a group (log out and back in to apply)
# note: membership in the docker group is effectively root access
sudo usermod -aG docker $USER
# change the owner and group of a file
sudo chown achilles:www-data file.txt
# change them recursively
sudo chown -R achilles:www-data /var/www/appFile Permission Modes and chmod
# Modifies the permissions for the user (owner) and others (o).
# The +x adds execute permissions.
chmod uo+x somefile.txt
# Modifies the permissions for the group (g).
# The -w removes write permissions for the group.
chmod g-w somefile.txt
# Octal modes: read = 4, write = 2, execute = 1, for user, group and others.
# rwxr-xr-x
chmod 755 script.sh
# rw-r--r--
chmod 644 file.txtProcesses
All running applications are processes.
A process is a running instance of one or more related tasks (threads). It is not the same thing as a program. A program or a command can start several simultaneous processes.
Some processes depend on each other and others are independent. A failed process may affect other processes running on the system.
In order to accomplish their task, they use system resources such as memory, CPU cycles and peripheral devices such as network cards, displays, etc. The operating system, particularly its kernel, plays a crucial role in distributing the right amount of these resources to each process while ensuring that the system as a whole is utilized efficiently.
Types of processes
Interactive Processes: These are user-initiated processes, started either via command line or graphical interface. Examples include ‘bash’ for command-line operations, ‘firefox’ as a web browser, ‘top’ for system monitoring, ‘Slack’ for communication etc.
Batch Processes: Automated and scheduled processes, running independently of user interaction and following a FIFO queue system. ‘updatedb’ and ‘ldconfig’ are common examples, often used for system maintenance tasks.
Daemons: Server processes that run in the background, often starting at system boot. They await system or user requests to provide services. Examples include ‘nginx’ for web services and ‘sshd’ for secure shell services.
Threads: Lightweight tasks running under a main process, sharing resources but individually managed. They enhance efficiency in complex applications like ‘dconf-service’ and ‘gnome-terminal-server’.
Kernel Threads: Essential system tasks managed by the kernel, such as ‘kthreadd’ and ‘migration’. These are not user-controlled and perform core system functions like resource management and process scheduling.
Process States
- Running (R): The process is either actively running on a CPU or is ready to run.
- Sleeping (S): The process is in a sleep state, waiting for an event or resource to become available.
- Uninterruptible Sleep (D): This is similar to Sleeping (S), but in this state, the process cannot be interrupted by signals. This usually happens when the process is waiting for I/O operations to complete.
- Stopped (T): The process has been stopped, either by a job control signal or because it is being traced.
- Zombie (Z): The process has completed its execution but still exists in the process table because the parent process has not yet read its exit status.
Identifiers associated with processes
- Process ID (PID): Unique Process ID number.
- Parent Process ID (PPID): Parent Process that initiated this process. Should the originating parent cease to function, the PPID will then point to an adoptive parent process; this role is typically assumed by init (‘systemd’), identified by a PPID of 1, or by a subreaper process.
- Thread ID (TID): For single-threaded processes, this identifier is identical to the PID. In contrast, within a multi-threaded process, although all threads share the same PID, each one possesses a distinct TID.
Process User and Group IDs
Each process is associated with user and group IDs, which determine the permissions and access levels of the process.
- RUID: Identifies the user who started the process.
- RGID: Identifies the group who started the process.
- EUID: Determines the access rights of the user.
- EGID: Determines the access rights of the group.
See running processes
# provides a detailed view of running processes
ps -l
# shows extensive details about all running processes
ps aux
# displays real-time information on system processes and resource usage
top
# an interactive and advanced version of top
htop
# displays the processes running on the system in the form of a tree diagram
pstreeProcess Priorities
The nice value of a process determines its priority level, with a lower nice value indicating a higher priority. Essential processes are usually given low niceness values, signifying high importance, whereas processes that can afford to wait are assigned higher values. A process with a higher niceness value, allows other processes to be executed first. In Linux, the niceness scale ranges from -20 for the highest priority to +19 for the lowest. This might seem counterintuitive, but this naming convention, where a ‘nicer’ process has a lower priority, is a legacy from the early days of UNIX.
Using nice and renice to Set Priorities
# detail view of all processes
ps axlf
# start a command with a nice value of 10
nice -n 10 command
# change the nice score of a process (lowering it needs sudo)
renice +5 3037Background and Foreground Processes
Typically, jobs run in the foreground as a standard. However, appending an & symbol to a command, such as updatedb &, will execute the job in the background instead.
To suspend a job running in the foreground, you can press CTRL-Z, which suspends it (SIGTSTP). To end the job entirely, use CTRL-C. If you have a suspended process, you can resume it in the background with the bg command, or bring a background process to the foreground using the fg command.
# start a process
gedit text.txt
# while on terminal hit CTRL+Z to suspend the process
CTRL+Z
# with jobs -l you can see what processes have been launched from this terminal
jobs -l
[1]+ 29554 Stopped gedit text.txt
# to run the process in the background use bg
bg
# to put the process in the foreground use fg
fgTerminating a Process
# find the process id
pgrep -a example_process
ps aux | grep example_process
# ask the process to stop (SIGTERM), so it can clean up
kill 1234
# force it (SIGKILL), only if it does not stop
kill -9 1234
# send SIGTERM to every process with that name
pkill example_processProcesses in containers
A container is a group of ordinary processes that share the host’s kernel and are isolated with namespaces and cgroups, which is why ps aux on the host lists them too.
Inside the container your app usually runs as PID 1, which gets no default signal handling, so it has to handle SIGTERM itself (or run behind a small init such as docker run --init), otherwise docker stop waits 10 seconds and sends SIGKILL.
Services (systemd)
Systemd is a system and service manager for Linux operating systems. It is the first process that starts when the Linux system boots and is the last process that stops when the system shuts down.
To manage systemd services use systemctl:
# check status of a service
systemctl status nginx
# start/stop/restart a service
sudo systemctl start nginx
sudo systemctl stop nginx
sudo systemctl restart nginx
# enable/disable a service on boot
sudo systemctl enable nginx
sudo systemctl disable nginx
# list services that failed
systemctl --failed
# reload unit files after editing them
sudo systemctl daemon-reloadLogs with journalctl
Systemd collects the logs of every service in the journal.
# follow the logs of a service (add sudo if you are not in the adm or systemd-journal group)
journalctl -u nginx -f
# logs of the last hour
journalctl -u nginx --since "1 hour ago"
# logs of the current boot, errors only
journalctl -b -p errScheduling Future Processes
Using cron
Cron is a time-based scheduling utility program.
# edit your own cron table
crontab -e
# list your cron table
crontab -l
# a line in it: minute hour day-of-month month day-of-week command
# run backup.sh every day at 02:30
30 2 * * * /home/achilles/backup.sh
# system-wide cron config file (cron table)
/etc/crontab
# anacron runs the jobs even if the system was down
# anacron config file
/etc/anacrontabSystemd timers
A systemd timer is a .timer unit that starts a matching .service unit on a schedule. Its output ends up in the journal, and with Persistent=true it catches up on runs it missed while the machine was off.
# list active timers and when they run next
systemctl list-timersThe At utility
at runs a command once at a given time. It is often not installed by default (package at).
# make a script executable
chmod +x testat.sh
at now + 1 minute -f testat.sh
# see if the job is queued
atq
# make sure the job actually ran
cat /tmp/datestamp
# interactively
at now + 1 minuteThe Sleep utility
# delays execution for a specific period
sleep 5 && gedit README.mdPackages
There are two types of package managers low and high level.
Low-level package managers handle the basic tasks of installing, removing, and managing individual packages. On the other hand, High-level package managers, like apt, simplify the process by automatically taking care of dependencies and providing easy access to a wide range of software through repositories, all within a user-friendly interface.
The examples below are for Debian and Ubuntu. Fedora and RHEL use rpm (low level) and dnf (high level) instead.
Low-level:
dpkg is a low-level package manager that is used to install, build, remove, and manage Debian packages.
# install a package
sudo dpkg -i package.deb
# remove a package (keep configuration files)
sudo dpkg -r package
# remove a package (remove configuration files)
sudo dpkg -P package
# list installed packages
dpkg -l
# list files installed by a package
dpkg -L packageHigh-level:
apt is a high-level package manager that is used to install, remove, and update software packages.
# update the package index
sudo apt update
# install a package
sudo apt install package1
# upgrade packages
sudo apt upgrade
# remove a package (keep configuration files)
sudo apt remove package1
# remove a package (remove configuration files)
sudo apt purge package1
# remove unused packages
sudo apt autoremove
# list installed packages
apt list --installed
# list files installed by a package
dpkg -L package1
# search for a package
apt search package1
# show package information
apt show package1Shell
Reading files
# concatenate files to standard output
cat file.txt
# displays the contents of a file one page at a time
less somefile
cat somefile | less
# display the beginning of a text file
head file.txt
head -n 5 somefile
head -5 somefile
# display the end of a file
tail file.txt
tail -15 somefile.log
# follow a file as it grows (logs)
tail -f somefile.log
# check a command manual
man catPipes
The pipe symbol | is used to pass the output of a program
as an input for another from left to right.
man head | headStandard file streams
i/o redirection
# change input stream
$ cat < file1.txt
# change output stream
$ cat file1.txt > file2.txt
# change output stream
$ cat file1.txt 1> file2.txt
# append to a file
$ cat file1.txt >> file2.txt
# redirect standard error to a file
$ command file1.txt 2> file2.txt
# redirect everything to a file
$ command file1.txt &> file2.txtDiscard output with /dev/null
ls -lR /tmp > /dev/nullgrep
Grep is used as a primary text searching tool. It scans files for specified patterns and can be used with regular expressions, as well as simple strings:
# search for a pattern in a file and print all matching lines
grep [pattern] somefile.txt
# print all lines that do not match the pattern
grep -v [pattern] somefile.txt
# print the lines that contain the numbers 0 through 9
grep '[0-9]' somefile.txt
# print the context of lines
grep -C 2 [pattern] somefile.txt
# search recursively in a directory
grep -r [pattern] directory/Locate applications
$ which cat
/usr/bin/cat
$ whereis cat
cat: /usr/bin/cat /usr/share/man/man1/cat.1.gz
$ type cat
cat is /usr/bin/catEnvironment variables
# display the value of a specific variable
echo $SHELL
# export a variable
export VARIABLE_NAME="value"
# add a variable permanently
echo 'export VARIABLE_NAME="value"' >> ~/.bashrc
# reload the bash config
source ~/.bashrcThe PATH Variable
The Path variable specifies directories where Linux shell searches for executable files to run commands without full paths.
You can add other things to Path variable:
export PATH=$HOME/bin:$PATHChange the command line prompt
# change the command line prompt to display the current working directory.
$ PS1="\w$ "
# change the command line prompt to display the current working directory and the time.
$ PS1="\w \t$ "History and shortcuts
# browse the list of executed commands
up/down keys
# execute previous command
!! (bang-bang)
# Search in command history
CTRL + R
# clears the Screen
CTRL + L
# temporarily halt output
CTRL + S
# continue temporarily halted output
CTRL + Q
# exit current shell (on an empty line)
CTRL + D
# suspend the current process
CTRL + Z
# interrupt the current process (SIGINT)
CTRL + C
# go to the beginning of the line
CTRL + A
# delete the word before the cursor
CTRL + W
# delete from the beginning of line to cursor position
CTRL + U
# go to the end of the line
CTRL + EThe VI editor
VI is a text editor for Linux that can be utilized when GUI is not available. On most distributions vi opens Vim.
It is more advanced than the Nano editor but it is more complex, too.
The key point to grasp is that VI operates in three distinct modes:
- Command mode
- The default mode, each key is an editor command.
- Insert mode
- Type “i” from command mode to switch to insert mode.
- Insert mode is used to edit the actual text in a file.
- Press “ESC” to exit insert mode and return to the Command mode.
- Line mode
- Type “:” to switch to the Line mode from the Command mode.
- Press “ESC” to exit Line mode and return to the Command mode.
# Edit some file
vi somefile.txt
# Recovery mode
vi -r somefile.txt
# In line mode
# write to a file
:w
# write and quit
:x or :wq
# quit
:q
# quit without saving
:q!
# Searching in VI
/pattern
# move to the next occurrence
n
# Move to previous occurrence
NNetworking
ip and ss (from the iproute2 package) replace the older ifconfig and netstat.
# show network interfaces and their addresses
ip a
# show the routing table
ip r
# show listening TCP and UDP ports and the processes using them
sudo ss -tulpn
# what is listening on port 8000?
sudo ss -tlnp | grep :8000
sudo lsof -i :8000
# check an HTTP endpoint (headers only, then verbose)
curl -I https://example.com
curl -v https://example.com- ethtool: ethtool is a Linux utility for network interface configuration, offering capabilities to view and change settings such as speed, duplex, and auto-negotiation.
- nmap: Scans open ports on a network. Useful for security analysis. Only scan networks and hosts you own or have permission to scan.
- tcpdump: Dumps network traffic.
- iptraf-ng: Monitors network traffic in text mode.
- mtr: Combines ping and traceroute.
- dig: Tests DNS lookups.
SSH
# create a key pair
ssh-keygen -t ed25519 -C "you@example.com"
# copy your public key to a server
ssh-copy-id user@server
# connect
ssh user@serverTo skip typing the details every time, add the server to ~/.ssh/config:
Host myserver
HostName 203.0.113.10
User deploy
IdentityFile ~/.ssh/id_ed25519Then ssh myserver is enough.