Virtual Machines¶
On this page
What virtualization actually does, the two hypervisor types, the VM lifecycle, snapshots vs backups, and the VM-vs-container decision.
Basics¶
A virtual machine (VM) is a software emulation of a physical computer. A hypervisor (Virtual Machine Monitor) sits between hardware and VMs, allocating CPU, memory, storage, and network so multiple isolated guest operating systems run on one host.
Type 1 vs Type 2 hypervisors¶
| Type 1 (bare-metal) | Type 2 (hosted) | |
|---|---|---|
| Runs on | Hardware directly | On top of a host OS |
| Examples | VMware ESXi, Microsoft Hyper-V, KVM/Proxmox, Xen | VirtualBox, VMware Workstation/Fusion, QEMU (user) |
| Use | Production, data centers, cloud | Desktops, dev/test, labs |
| Overhead | Lower | Higher |
KVM is interesting: it's a Linux kernel module that turns Linux into a Type 1 hypervisor. Proxmox VE and most clouds build on KVM.
How it works¶
Modern CPUs provide hardware-assisted virtualization (Intel VT-x, AMD AMD-V) so guests run near-native speed. Nested page tables (EPT/NPT) virtualize memory. Paravirtualized drivers (virtio) give fast disk/network I/O. IOMMU (VT-d/AMD-Vi) enables passing a physical device (GPU, NIC) directly to a VM (PCI passthrough).
Key resources¶
- vCPU — virtual CPUs mapped onto physical cores/threads (can oversubscribe).
- RAM — assigned per VM; ballooning reclaims unused memory.
- Virtual disk — a file/volume (
qcow2,vmdk,raw, LVM, ZFS). - Virtual NIC — attached to a bridge/virtual switch.
VM lifecycle¶
Provision → Configure (OS + apps) → Run → Snapshot → Migrate → Backup → Decommission
- Snapshot — point-in-time state (disk ± RAM) you can roll back to. Not a backup.
- Live migration — move a running VM between hosts with minimal downtime (shared/replicated storage).
- Cloning — copy a VM; templates are golden images for fast provisioning.
- Thin vs thick provisioning — allocate disk on demand vs upfront.
Cheatsheet¶
KVM / libvirt (virsh)¶
virsh list --all # all VMs (domains)
virsh start myvm
virsh shutdown myvm # graceful (ACPI)
virsh destroy myvm # force off (not delete)
virsh autostart myvm
virsh dominfo myvm
virsh console myvm
virsh snapshot-create-as myvm snap1 "before upgrade"
virsh snapshot-list myvm
virsh snapshot-revert myvm snap1
virsh migrate --live myvm qemu+ssh://host2/system
virt-install ... # create a new VM
QEMU (quick disk + boot)¶
qemu-img create -f qcow2 disk.qcow2 40G
qemu-img info disk.qcow2
qemu-img convert -O qcow2 in.vmdk out.qcow2
qemu-img resize disk.qcow2 +10G
Vagrant (dev VMs as code)¶
Check virtualization support¶
egrep -c '(vmx|svm)' /proc/cpuinfo # >0 means VT-x/AMD-V present
kvm-ok # Ubuntu helper
lscpu | grep Virtualization
Thumb Rules¶
Rules of thumb
- A snapshot is not a backup. Snapshots depend on the original disk; a backup is independent.
- Don't keep snapshots for long. Snapshot chains grow, hurt performance, and can corrupt — delete after the change is verified.
- Don't oversubscribe RAM the way you oversubscribe CPU. CPU overcommit is usually fine; memory overcommit causes swapping and instability.
- One workload's noisy I/O hurts neighbors. Watch disk latency on shared storage.
- Use templates/golden images for consistency instead of hand-building each VM.
- Right-size, then scale. Start small; it's easier to add vCPU/RAM than to reclaim it.
- Install guest agents (qemu-guest-agent / VMware Tools) for clean shutdown, time sync, and balloon.
Use Cases¶
- Server consolidation — many workloads on fewer physical machines.
- Isolation & multi-tenancy — strong boundaries between guests (stronger than containers).
- Legacy & mixed OS — run Windows and Linux side by side, or old OSes.
- Dev/test & labs — disposable environments, snapshots before risky changes.
- Disaster recovery — replicate VMs to a second site.
- Cloud IaaS — EC2/GCE/Azure instances are VMs under the hood.
Common Issues¶
VM won't boot after enabling virtualization features
Virtualization (VT-x/AMD-V) may be disabled in BIOS/UEFI, or nested virtualization
isn't enabled. Check egrep -c '(vmx|svm)' /proc/cpuinfo.
Poor disk/network performance
You're likely using emulated (IDE/e1000) devices instead of virtio paravirtualized drivers. Switch to virtio and install guest drivers.
Snapshot grew huge / VM slow
A long-lived snapshot's delta file keeps growing. Consolidate/delete the snapshot. Never run production on a days-old snapshot.
Time drift inside the guest
Install the guest agent and sync via host or NTP/chrony. Drift breaks TLS, Kerberos, logs.
Live migration fails
Usually mismatched CPU models/flags between hosts, or storage not shared/replicated. Use a common CPU baseline and shared storage.
Out of memory / heavy swapping on the host
RAM overcommitted. Reduce assignments, enable ballooning, or add memory. Memory is the resource you should not aggressively oversubscribe.
Best Practices¶
- Enable virtio + guest agents for performance and clean lifecycle operations.
- Separate storage tiers (fast SSD/NVMe for OS/DB, bulk for archives).
- Template golden images and provision from them (Packer to build, Terraform to deploy).
- Back up independently of snapshots; test restores regularly.
- Patch the hypervisor — it's the most security-critical layer.
- Reserve resources for the host; never allocate 100% of RAM to guests.
- Monitor host and guest (CPU ready time, memory ballooning, disk latency).
- Use PCI/GPU passthrough only when you need it; it complicates migration.
Official Sources¶
- KVM — https://linux-kvm.org/page/Documents
- libvirt /
virsh— https://libvirt.org/docs.html - QEMU documentation — https://www.qemu.org/docs/master/
- VMware vSphere / ESXi docs — https://techdocs.broadcom.com/us/en/vmware-cis/vsphere.html
- Microsoft Hyper-V — https://learn.microsoft.com/en-us/windows-server/virtualization/hyper-v/hyper-v-on-windows-server
- HashiCorp Vagrant — https://developer.hashicorp.com/vagrant/docs
- HashiCorp Packer (golden images) — https://developer.hashicorp.com/packer/docs