Skip to content

Ansible

On this page

Ansible's agentless, push-based automation: inventory, playbooks, roles, modules, variables, and the idempotency model that makes reruns safe.

Basics

Ansible (Red Hat) automates configuration, deployment, and orchestration. Its defining traits:

  • Agentless — connects over SSH (Linux) or WinRM (Windows); nothing to install on targets except Python.
  • Push-based — you run from a control node and push changes out.
  • Declarative + idempotent — modules describe desired state and only change what's needed.
  • YAML — playbooks are human-readable YAML, not code.

Core concepts

Concept What it is
Control node Where you run ansible (any Linux/macOS with Python; not Windows).
Managed node The target host (needs SSH + Python).
Inventory List of hosts/groups (INI, YAML, or dynamic from cloud APIs).
Module A unit of work (apt, copy, service, user). Runs on the target, returns JSON.
Task A single call to a module.
Play Maps a group of hosts to tasks.
Playbook An ordered list of plays (a YAML file).
Role Reusable, structured bundle of tasks/handlers/templates/vars/defaults.
Handler A task triggered by notify (e.g. restart service), runs once at end.
Collection Distributable package of roles/modules/plugins (via Ansible Galaxy).
Facts System data gathered automatically (via setup).

Execution model

  1. Parse inventory and variables.
  2. For each play, gather facts (optional).
  3. Run tasks in order, host by host (configurable forks/strategy).
  4. Modules execute on the target and report ok / changed / failed.
  5. Handlers fire if notified.

Cheatsheet

Inventory (inventory.ini)

[web]
web1 ansible_host=10.0.0.11
web2 ansible_host=10.0.0.12

[db]
db1 ansible_host=10.0.0.21

[prod:children]
web
db

[all:vars]
ansible_user=ubuntu
ansible_ssh_private_key_file=~/.ssh/id_ed25519

Ad-hoc commands

ansible all -i inventory.ini -m ping
ansible web -m apt -a "name=nginx state=present" --become
ansible all -m shell -a "uptime"
ansible all -m setup            # dump facts

A playbook (site.yml)

- name: Configure web servers
  hosts: web
  become: true
  vars:
    doc_root: /var/www/site
  tasks:
    - name: Install nginx
      ansible.builtin.apt:
        name: nginx
        state: present
        update_cache: true

    - name: Deploy config
      ansible.builtin.template:
        src: nginx.conf.j2
        dest: /etc/nginx/nginx.conf
        mode: "0644"
      notify: Reload nginx

    - name: Ensure nginx running
      ansible.builtin.service:
        name: nginx
        state: started
        enabled: true

  handlers:
    - name: Reload nginx
      ansible.builtin.service:
        name: nginx
        state: reloaded

Run, check, target

ansible-playbook -i inventory.ini site.yml
ansible-playbook site.yml --check --diff     # dry run + show changes
ansible-playbook site.yml --limit web1       # one host
ansible-playbook site.yml --tags deploy
ansible-playbook site.yml -e "doc_root=/srv" # extra vars (highest precedence)

Roles, Galaxy, Vault

ansible-galaxy init roles/web                # scaffold a role
ansible-galaxy collection install community.general
ansible-vault create secrets.yml            # encrypted vars
ansible-vault edit secrets.yml
ansible-playbook site.yml --ask-vault-pass
ansible-lint                                # lint playbooks/roles

Role structure

roles/web/
├── defaults/main.yml   # lowest-precedence vars
├── vars/main.yml       # higher-precedence vars
├── tasks/main.yml
├── handlers/main.yml
├── templates/
├── files/
└── meta/main.yml       # dependencies

Thumb Rules

Rules of thumb

  • Use modules, avoid raw shell/command. They're idempotent and report changed correctly. If you must, guard with creates/when.
  • Name every task. Output and --start-at-task depend on it.
  • Run --check --diff first on anything touching production.
  • Roles for reuse, playbooks for orchestration. Keep playbooks thin.
  • Vault all secrets. Never commit plaintext credentials.
  • Pin collection/role versions in requirements.yml for reproducible runs.
  • Idempotency is the test: a second run should report changed=0.
  • Prefer become per play/task over running as root everywhere.
  • Use FQCN (ansible.builtin.copy) to avoid module name clashes.

Use Cases

  • Configuration management — packages, files, services, users across fleets.
  • Application deployment — rolling deploys with serial and health checks.
  • Provisioning + orchestration — call cloud modules (AWS/Azure/GCP), then configure.
  • One-off operations — patch waves, cert rotation, emergency fixes via ad-hoc.
  • Network & appliance automation — many vendor collections.
  • CI/CD glue — invoked from pipelines for environment setup.

Common Issues

“Permission denied” / sudo password required

Missing become: true or no privilege escalation password. Use --become and --ask-become-pass, or configure passwordless sudo for the ansible user.

Host unreachable / SSH errors

Wrong ansible_user, key, or host key prompt. Test with ansible host -m ping -vvv. For first contact, set host_key_checking appropriately or pre-seed known_hosts.

Task reports changed every run

A command/shell task with no guard. Add creates:/removes:/when: or switch to a stateful module. This breaks idempotency.

Python interpreter not found on target

Set ansible_python_interpreter (e.g. /usr/bin/python3) or install Python on the node.

Variable precedence confusion

Same var defined in multiple places resolves by Ansible's precedence order (extra-vars win, defaults lose). Keep variable sources minimal; debug with the debug module.

Handlers didn't run

Handlers only run when notified and the play reaches the end without a fatal error, unless --force-handlers. A failed earlier task can skip them.

Best Practices

  • Structure with roles; keep a clear group_vars/, host_vars/ layout.
  • Encrypt secrets with Ansible Vault; integrate an external secret store for scale.
  • Lint (ansible-lint) and dry-run in CI; use molecule for role testing.
  • Pin versions of collections/roles in requirements.yml.
  • Use tags and --limit for safe, targeted runs.
  • Deploy gradually with serial: and health checks to avoid taking everything down at once.
  • Make runs idempotent and verify with changed=0 on reruns.
  • Keep inventory in source control (or use dynamic inventory for cloud).

Official Sources