- Python 99.6%
- Dockerfile 0.2%
- PLpgSQL 0.1%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .forgejo/workflows | ||
| airflow | ||
| docs | ||
| k8s-argocd | ||
| k8s-datajob | ||
| migrations | ||
| ops | ||
| tests | ||
| .dockerignore | ||
| .env.example | ||
| .gitattributes | ||
| .gitignore | ||
| docker-compose.cloud.yaml | ||
| docker-compose.dev.yaml | ||
| docker-compose.yaml | ||
| Dockerfile | ||
| Dockerfile.base | ||
| README.md | ||
| requirements-app.txt | ||
| requirements.txt | ||
DATAJOB Airflow
The complete Airflow runtime: images, DAGs, scripts, SQL, migrations, bootstrap, tests and deployment configuration.
This repository is the canonical source for Airflow runtime code. It is not duplicated anywhere else.
Images
Two Dockerfiles, built in order.
apache/airflow:3.2.2
|
v Dockerfile.base
datajob/airflow-base:3.2.2-<n> dependencies only, no application code
|
v Dockerfile
tanss-airflow:<git-sha> + DAGs, scripts, SQL, migrations, bootstrap
Dockerfile.base builds the reusable dependency image: apache/airflow:3.2.2
plus the platform Python dependencies, installed under Airflow's published
constraints.
Dockerfile builds the final application image on top of it. That image is
what every Airflow component runs — API server, scheduler, DAG processor,
triggerer, both Celery workers, the bootstrap job and the migration job. The
component is chosen by the container command, never by the image.
docker build -f Dockerfile.base -t datajob/airflow-base:3.2.2-1 .
docker build -f Dockerfile -t tanss-airflow:$(git rev-parse --short HEAD) .
What is in the application image
| Container path | Source | Used by |
|---|---|---|
/opt/airflow/dags |
airflow/dags/ |
DAG processor discovers these at startup |
/opt/airflow/scripts |
airflow/scripts/ |
python /opt/airflow/scripts/<x>.py |
/opt/airflow/sql |
airflow/sql/ |
validation and maintenance SQL |
/opt/airflow/migrations |
migrations/ |
run_migrations.py via MIGRATIONS_DIR |
Bootstrap (airflow/scripts/bootstrap_airflow.py) is baked in with the rest.
There is no separate step to register DAGs: the files are in the image, and
the DAG processor picks them up.
.dockerignore is an allow-list, so a new secret or data file added to the
repository cannot silently reach an image layer.
Databases and Redis
Owned by this repository, internal to the Airflow stack:
- Airflow metadata PostgreSQL — Airflow's own state
- Redis — the Celery broker
Neither is reachable from the TANSS stack.
Owned by tanss_ticket_analysis, shared:
tanss-postgres— the business and ticket data. Airflow writes to it and Metabase reads from it, so it deliberately lives outside this repository.
Airflow reaches it over the shared tanss_data_net Docker network by DNS
name. Because it belongs to a separate Compose project there is no
depends_on for it: tanss-db-migrate polls until it is reachable.
That repository also keeps Metabase and its models, embedding_api, the
Caddy configuration and the backup tooling for the business databases. The
Airflow metadata database is backed up here, by
ops/backup_airflow_metadata.sh.
Create the shared networks once, on either host:
docker network create tanss_data_net
docker network create tanss_edge_net
Dependencies
apache/airflow:3.2.2 already ships psycopg2-binary, python-dotenv,
requests, httpx, pandas, numpy and the celery / fab / postgres /
standard providers. Do not re-add them.
requirements.txt holds platform-wide dependencies for the base image.
requirements-app.txt holds anything specific to this application.
Never install in the Airflow images
sentence-transformers, torch, triton, and any nvidia-* wheel. They
pull the CUDA/NVIDIA stack, which makes every Airflow container slow to
start. Only the dedicated embedding API service needs them, and production
runs EMBEDDING_PROVIDER=api so Airflow talks to it over HTTP instead.
Both Dockerfiles fail the build if these appear, matching nvidia- by prefix
so a newer CUDA generation cannot slip past a pinned package list.
Never use _PIP_ADDITIONAL_REQUIREMENTS
It installs packages on every container start, and Airflow documents it as a development-only convenience. Setting it once left the default Celery worker unable to finish starting while tasks queued behind it. Dependencies belong in the image at build time; adding one means editing the requirements file and rebuilding.
Bootstrap
airflow-bootstrap runs after airflow-init and before the long-running
components. It is idempotent and safe to run on every deployment:
- verify the metadata database is reachable
- create missing Airflow pools
- verify required secrets are configured, reported by name only
- verify every DAG parses
It never triggers, unpauses or schedules a DAG, never creates a secret, and never modifies a pool that already exists — a hand-tuned slot count survives a deployment.
Kubernetes
Docker Compose is transitional. The image itself carries no Compose
assumptions: no bind mounts, no depends_on, no host .env, no runtime pip,
and nothing written to the container filesystem beyond /opt/airflow/logs.
The same immutable image is intended for Kubernetes Deployments and Jobs, with Redis, both PostgreSQL instances and all credentials supplied through service DNS, ConfigMaps and Secrets. Hostnames and credentials are never baked into the image.
See docs/airflow_image_and_deployment.md for the resource mapping and the
infrastructure decisions still open.
Tagging
datajob/airflow-base:<airflow-version>-<n> bump <n> when dependencies change
tanss-airflow:<git-sha> one tag per deployed commit
Never deploy latest. Every Airflow service in a release runs the same tag.