DevOps Elastic Hayway
Document

SUBSCRIBE TO GET FULL ACCESS TO THE E-BOOKS FOR FREE 🎁SUBSCRIBE NOW

Professional Dropdown with Icon

SUBSCRIBE NOW TO GET FREE ACCESS TO EBOOKS

DockerLesson 16 / 227 min readUpdated September 11, 2026

Docker Swarm Tutorial: Build a Cluster, Deploy a Stack, and Compare It with Kubernetes

Docker Swarm is the clustering mode built into Docker Engine. With two commands you turn a set of Docker hosts into a cluster that schedules containers, load-balances traffic and restarts failed replicas. It is far simpler than Kubernetes and still maintained in 2026 (Mirantis ships it as part of MKE), which makes it a legitimate choice for small teams and edge deployments and an excellent first step before Kubernetes. This tutorial builds a three-node swarm, deploys a stack from a Compose file, scales and updates it, and closes with a comparison that helps you choose between Swarm and Kubernetes.

Prerequisites: three Linux hosts with Docker Engine 27+ installed (Docker installation), open ports TCP 2377 (cluster management), TCP/UDP 7946 (node gossip) and UDP 4789 (overlay network). Knowledge of Docker Compose helps.

Swarm concepts in one minute

  • Node: a Docker host in the swarm; managers hold the Raft-replicated cluster state and schedule work, workers run containers. Use an odd number of managers (1, 3 or 5).
  • Service: the desired state of a container (image, replicas, ports, networks). Swarm creates tasks (one container each) to satisfy it.
  • Stack: a group of services described in a Compose file and deployed together.
  • Overlay network: a virtual network spanning all nodes; services on it reach each other by name.
  • Routing mesh: any published port is reachable on every node and load-balanced to the tasks.

Step 1 – Initialise the swarm

# On the first node (manager)
docker swarm init --advertise-addr 192.168.56.10
# Swarm initialized: current node (x1y2z3) is now a manager.
# To add a worker to this swarm, run the following command:
#   docker swarm join --token SWMTKN-1-<worker-token> 192.168.56.10:2377

docker swarm join-token worker     # print again later
docker swarm join-token manager    # token for additional managers

Step 2 – Join the workers

# On each worker
docker swarm join --token SWMTKN-1-<worker-token> 192.168.56.10:2377

# Back on the manager
docker node ls
# ID        HOSTNAME   STATUS   AVAILABILITY   MANAGER STATUS   ENGINE VERSION
# x1y2z3 *  manager1   Ready    Active         Leader           27.5.1
# a4b5c6    worker1    Ready    Active                          27.5.1
# d7e8f9    worker2    Ready    Active                          27.5.1

# Optional: keep application tasks off the manager
docker node update --availability drain manager1
docker node update --label-add tier=frontend worker1

Step 3 – Your first service

docker service create --name web --replicas 3 --publish published=8080,target=80 nginx:1.27-alpine

docker service ls
docker service ps web            # which node runs each task
curl -s http://192.168.56.11:8080 | head -3     # any node answers (routing mesh)

# Scale and inspect
docker service scale web=6
docker service logs -f web
docker service inspect web --pretty

Kill a container (docker rm -f on a worker) or shut a worker down: within seconds Swarm reschedules the missing tasks elsewhere. That self-healing is the core promise of any orchestrator.

Step 4 – Deploy a stack from a Compose file

Stacks reuse the Compose format with a deploy: section that Swarm understands (and plain docker compose ignores). This example runs a web tier, a Redis cache and a visualizer, with a secret and a placement constraint.

# stack.yml
services:
  web:
    image: ghcr.io/myorg/vote:1.4.2
    ports:
      - target: 80
        published: 8080
        mode: ingress                 # routing mesh (default); use "host" to bypass it
    networks: [frontend, backend]
    secrets: [db_password]
    environment:
      REDIS_HOST: redis
      DB_PASSWORD_FILE: /run/secrets/db_password
    deploy:
      replicas: 4
      placement:
        constraints:
          - node.role == worker
          - node.labels.tier == frontend
      update_config:
        parallelism: 1
        delay: 10s
        order: start-first            # start the new task before stopping the old one
        failure_action: rollback
      rollback_config:
        parallelism: 2
      restart_policy:
        condition: on-failure
        max_attempts: 3
      resources:
        limits: { cpus: "0.50", memory: 256M }
        reservations: { cpus: "0.10", memory: 64M }
    healthcheck:
      test: ["CMD", "wget", "-qO-", "http://localhost/healthz"]
      interval: 15s
      timeout: 3s
      retries: 3

  redis:
    image: redis:7.4-alpine
    networks: [backend]
    volumes:
      - redis-data:/data
    deploy:
      replicas: 1
      placement:
        constraints: [node.hostname == worker2]   # stateful → pin to the node with the volume

  visualizer:
    image: dockersamples/visualizer:stable
    ports: ["8081:8080"]
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock
    deploy:
      placement:
        constraints: [node.role == manager]

networks:
  frontend:
    driver: overlay
  backend:
    driver: overlay
    internal: true                    # no outbound internet from this network

volumes:
  redis-data:

secrets:
  db_password:
    external: true                    # created with docker secret create
# Create the secret from a file (never from the command line history)
docker secret create db_password ./db_password.txt

docker stack deploy -c stack.yml vote
docker stack services vote
docker stack ps vote --no-trunc

# Open http://<any-node-ip>:8081 to see the task distribution

Step 5 – Rolling update and rollback

# Update the image: one task at a time, 10 s apart, start-first
docker service update --image ghcr.io/myorg/vote:1.5.0 vote_web
watch docker service ps vote_web --filter desired-state=running

# Something wrong? Roll back to the previous spec
docker service rollback vote_web

# Or redeploy the stack after editing stack.yml (idempotent)
docker stack deploy -c stack.yml vote

Step 6 – Day-2 operations

# Add a manager for high availability (3 managers tolerate 1 failure)
docker swarm join-token manager        # run on a new node
docker node promote worker1            # or promote an existing worker

# Maintenance on a node: drain, patch, reboot, activate
docker node update --availability drain worker2
docker node update --availability active worker2

# Backup the Raft state (on a manager, stop Docker first for a consistent copy)
sudo systemctl stop docker && sudo tar czf swarm-backup.tgz /var/lib/docker/swarm && sudo systemctl start docker

# Rotate the join tokens and the CA if a token leaked
docker swarm join-token --rotate worker
docker swarm ca --rotate

# Remove the stack / leave the swarm
docker stack rm vote
docker swarm leave --force            # on a manager

Docker Swarm vs Kubernetes (2026)

Docker SwarmKubernetes
SetupTwo commands; built into Dockerkubeadm, k3s, or a managed service (EKS/AKS/GKE)
Learning curveDays; reuses Compose filesWeeks to months; dozens of object types
SchedulingSpread/constraints/labelsAffinity, taints, priorities, custom schedulers
NetworkingOverlay + routing mesh, built inCNI plugin of your choice, NetworkPolicy, Ingress/Gateway API
StorageVolume plugins, mostly node-localCSI drivers, PersistentVolumes, dynamic provisioning
AutoscalingNone built in (manual scale)HPA, VPA, Cluster Autoscaler/Karpenter
Secrets/configSecrets and configs, encrypted at rest in RaftSecrets (base64, encrypt with KMS), ConfigMaps, external operators
EcosystemSmall: Portainer, Traefik, SwarmpitHuge: Helm, ArgoCD, Prometheus Operator, service meshes, CNCF landscape
Managed cloud offeringNone from the big threeEvery cloud
Job marketNicheStandard skill for DevOps roles
Sweet spot1–20 nodes, small team, edge/IoT, simple web stacksAnything larger, multi-team, cloud-native platforms

How to decide

  • Choose Swarm if you already run Docker Compose in production, have a handful of hosts, no dedicated platform team, and no need for autoscaling or advanced storage. It gets you HA, rolling updates and secrets for almost zero learning cost.
  • Choose Kubernetes if you need autoscaling, multi-tenant isolation, a managed control plane, GitOps tooling, or if hiring and career growth matter. For small clusters, k3s gives Kubernetes with Swarm-like simplicity of installation.
  • Middle path: keep Compose files as the source of truth, deploy to Swarm today, and migrate later with kompose or by rewriting manifests; the concepts (services, replicas, networks, secrets, rolling updates) map one-to-one.

Troubleshooting

  • Tasks stuck in “Pending” / “no suitable node” – placement constraint matches no node, or the image cannot be pulled on workers (private registry: deploy with --with-registry-auth).
  • Published port works on one node only – UDP 4789 or TCP/UDP 7946 blocked between nodes; the routing mesh needs them.
  • Cluster lost quorum – with 3 managers and 2 down, the swarm is read-only; recover with docker swarm init --force-new-cluster on the surviving manager.
  • Service keeps restarting – check docker service ps --no-trunc for the error, then the healthcheck; start-first updates need enough spare resources for one extra task.

Key takeaways

  • docker swarm init + docker swarm join create a cluster; odd number of managers for quorum.
  • Services declare desired state; stacks group services from a Compose file with a deploy: section.
  • Rolling updates, rollback, secrets, overlay networks and the routing mesh are built in.
  • Swarm wins on simplicity for small clusters; Kubernetes wins on ecosystem, scale and career value.

Next tutorial

Next: Kubernetes learning path starting with installing a cluster with kubeadm. Official docs: Swarm mode overview, Compose deploy specification.

Retour parcours Docker — hub de la série et leçons sœurs.

Share your love

Leave a Reply

Your email address will not be published. Required fields are marked *