Overview

The first time someone explained operators to me, they said "it's a controller that manages a custom resource." That's accurate and useless. The second explanation was "you teach Kubernetes how to manage your application the same way it manages Deployments." That one helped.

If you've written a Helm chart, you know the problem operators solve: Helm installs a set of resources once and walks away. It doesn't reconcile drift, doesn't handle upgrades based on runtime state, doesn't restart a failed node in a Database cluster. An operator does.

CRDs vs built-in resources

Built-in (Deployment, Service)Custom Resource (CRD)
Defined inKubernetes sourceYour YAML
API endpoint/apis/apps/v1/apis/yourgroup.io/v1
ControllerShips with k8sYou write it
Storageetcdetcd
kubectl supportNativeNative, once CRD is applied

A CRD adds a new resource type. Once applied, kubectl get postgresclusters works like any other resource. The API server stores your custom objects in etcd, validates them against the CRD's schema, and hands them to whoever is watching.

The controller — the operator — is what makes the resource actually do something. Without a controller, a CRD is just a schema for storing YAML.

A minimal CRD

apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: websites.example.com
spec:
  group: example.com
  scope: Namespaced
  names:
    plural: websites
    singular: website
    kind: Website
    shortNames:
      - web
  versions:
    - name: v1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              required: [image, replicas]
              properties:
                image:
                  type: string
                replicas:
                  type: integer
                  minimum: 0
                  maximum: 10
                domain:
                  type: string
            status:
              type: object
              properties:
                readyReplicas:
                  type: integer
                url:
                  type: string
      subresources:
        status: {}
      additionalPrinterColumns:
        - name: Image
          type: string
          jsonPath: .spec.image
        - name: Replicas
          type: integer
          jsonPath: .spec.replicas
        - name: Age
          type: date
          jsonPath: .metadata.creationTimestamp

Two details worth calling out.

The subresources.status block. Adding this creates a separate /status endpoint so the controller can update status without touching spec. Without it, you get conflict errors when a user edits the resource while your controller is also writing to it.

additionalPrinterColumns. This is what makes kubectl get websites show useful columns instead of just NAME and AGE. Every operator CRD should have these — otherwise users have to describe each resource to see anything.

The controller pattern

A controller runs a reconciliation loop:

  1. Watch for changes to resources of the type you care about
  2. When something changes, compute the desired state
  3. Compare to the actual state
  4. Take action to close the gap
  5. Update status
  6. Repeat

The critical property is that reconciliation is idempotent. Running it twice should produce the same result. The controller doesn't remember "I already created this Deployment" — it checks whether the Deployment exists and matches the spec, and creates or updates accordingly.

func (r *WebsiteReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
    log := log.FromContext(ctx)

    // 1. Fetch the custom resource
    var website examplev1.Website
    if err := r.Get(ctx, req.NamespacedName, &website); err != nil {
        if apierrors.IsNotFound(err) {
            return ctrl.Result{}, nil
        }
        return ctrl.Result{}, err
    }

    // 2. Build the desired Deployment
    desired := &appsv1.Deployment{
        ObjectMeta: metav1.ObjectMeta{
            Name:      website.Name,
            Namespace: website.Namespace,
            Labels:    labelsForWebsite(&website),
        },
        Spec: appsv1.DeploymentSpec{
            Replicas: &website.Spec.Replicas,
            Selector: &metav1.LabelSelector{
                MatchLabels: labelsForWebsite(&website),
            },
            Template: corev1.PodTemplateSpec{
                ObjectMeta: metav1.ObjectMeta{
                    Labels: labelsForWebsite(&website),
                },
                Spec: corev1.PodSpec{
                    Containers: []corev1.Container{{
                        Name:  "web",
                        Image: website.Spec.Image,
                        Ports: []corev1.ContainerPort{{ContainerPort: 8080}},
                    }},
                },
            },
        },
    }

    // 3. Create or update
    var existing appsv1.Deployment
    err := r.Get(ctx, types.NamespacedName{Name: website.Name, Namespace: website.Namespace}, &existing)
    if apierrors.IsNotFound(err) {
        if err := r.Create(ctx, desired); err != nil {
            return ctrl.Result{}, err
        }
    } else if err != nil {
        return ctrl.Result{}, err
    } else if !deploymentMatches(&existing, desired) {
        existing.Spec = desired.Spec
        if err := r.Update(ctx, &existing); err != nil {
            return ctrl.Result{}, err
        }
    }

    // 4. Update status
    website.Status.ReadyReplicas = existing.Status.ReadyReplicas
    website.Status.URL = fmt.Sprintf("https://%s", website.Spec.Domain)
    if err := r.Status().Update(ctx, &website); err != nil {
        return ctrl.Result{}, err
    }

    return ctrl.Result{}, nil
}

That's the shape of it. Real operators do more — handle finalizers, watch multiple resource types, requeue with backoff — but the loop structure is the same.

Owner references: the garbage collection mechanism

if err := ctrl.SetControllerReference(&website, desired, r.Scheme); err != nil {
    return ctrl.Result{}, err
}

Setting the controller reference tells Kubernetes that the Deployment is owned by the Website. When the Website is deleted, the Deployment is garbage collected automatically. Without this, deleting a Website leaves orphaned resources that you have to clean up manually.

This is the single most forgotten line in operator code, and the consequence is that uninstalling your CRD leaves a mess.

Finalizers: cleanup on delete

Sometimes deleting a resource needs to do something first — delete an S3 bucket, deregister from a DNS server, wait for a database to shut down cleanly.

const finalizerName = "example.com/cleanup"

func (r *WebsiteReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
    var website examplev1.Website
    if err := r.Get(ctx, req.NamespacedName, &website); err != nil {
        return ctrl.Result{}, client.IgnoreNotFound(err)
    }

    if website.DeletionTimestamp.IsZero() {
        // Not being deleted — add the finalizer if missing
        if !controllerutil.ContainsFinalizer(&website, finalizerName) {
            controllerutil.AddFinalizer(&website, finalizerName)
            if err := r.Update(ctx, &website); err != nil {
                return ctrl.Result{}, err
            }
        }
    } else {
        // Being deleted — do cleanup
        if controllerutil.ContainsFinalizer(&website, finalizerName) {
            if err := r.cleanupExternalResources(ctx, &website); err != nil {
                return ctrl.Result{}, err  // retry on next reconcile
            }
            controllerutil.RemoveFinalizer(&website, finalizerName)
            if err := r.Update(ctx, &website); err != nil {
                return ctrl.Result{}, err
            }
        }
        return ctrl.Result{}, nil
    }

    // normal reconciliation continues here
}

The trap with finalizers: if the cleanup fails and never succeeds, the resource is stuck in a terminating state forever. Users can't delete it. Add a timeout or a manual escape hatch.

# Get stuck resources unstuck
kubectl patch website my-site -p '{"metadata":{"finalizers":[]}}' --type=merge

Kubebuilder vs Operator SDK

FrameworkLanguageUse when
KubebuilderGoStandard choice for Go operators
Operator SDKGo, Ansible, HelmYou want Ansible or Helm-based operators
KOPFPythonPython team, simpler operators
kube-rsRustRust team

Kubebuilder is the reference implementation — Operator SDK wraps it. If you're writing Go, use Kubebuilder directly. The scaffolding is:

kubebuilder init --domain example.com --repo GitHub.com/you/website-operator
kubebuilder create api --group example --version v1 --kind Website
make manifests
make install
make run

You get a working controller with the reconcile loop stubbed out, RBAC manifests generated from kubebuilder markers, and a test suite.

RBAC: the part that bites in production

// +kubebuilder:rbac:groups=example.com,resources=websites,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=example.com,resources=websites/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete

func (r *WebsiteReconciler) Reconcile(...) { ... }

These markers generate a ClusterRole when you run make manifests. Every resource your controller touches needs a corresponding marker. Forgetting one means the controller works in dev (where you run with admin credentials) and fails in production (where the service account doesn't have permission).

The most common mistake: not including status as a separate resource. websites and websites/status are different RBAC entries.

Testing operators

The envtest package runs a real API server and etcd without a full cluster. It's fast enough for integration tests.

func TestWebsiteReconciler(t *testing.T) {
    ctx := context.Background()

    // Create a Website
    website := &examplev1.Website{
        ObjectMeta: metav1.ObjectMeta{Name: "test", Namespace: "default"},
        Spec: examplev1.WebsiteSpec{
            Image:    "nginx:1.27",
            Replicas: 3,
        },
    }
    Expect(k8sClient.Create(ctx, website)).To(Succeed())

    // Reconcile
    _, err := reconciler.Reconcile(ctx, reconcile.Request{
        NamespacedName: types.NamespacedName{Name: "test", Namespace: "default"},
    })
    Expect(err).NotTo(HaveOccurred())

    // Assert the Deployment was created
    var dep appsv1.Deployment
    Eventually(func() error {
        return k8sClient.Get(ctx, types.NamespacedName{Name: "test", Namespace: "default"}, &dep)
    }).Should(Succeed())

    Expect(*dep.Spec.Replicas).To(Equal(int32(3)))
}

This pattern — create, reconcile, assert — catches most controller bugs before they hit a real cluster.

When not to write an operator

  • Your application has no runtime state to manage. A Helm chart is enough.
  • The logic is a one-time setup. Jobs and init containers are simpler.
  • You're just wrapping existing resources. If it doesn't need custom logic to reconcile, don't create a CRD.
  • You can use an existing operator. CloudNativePG, Percona, Zalando, and dozens of others already do what you'd be building.

The sweet spot is stateful applications — databases, message brokers, caches — where the correct behavior depends on runtime state that a static manifest can't express. For everything else, the complexity rarely pays off.