5 Commits

Author SHA1 Message Date
claude bb37314f31 chore: drop the throwaway comment used to force a rebuild
Build and Push Docker Images / build (push) Successful in 35s
Build and Push Docker Images / smoke (push) Successful in 0s
It existed only to make the image content differ so a deploy could be
timed. A note about a one-off measurement does not belong in a build file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WpTYCBLH58XzrM7n3xPJ5N
2026-08-24 11:31:23 +00:00
claude 4708b04482 chore: note the first myAi deploy with the Watchtower trigger configured
Build and Push Docker Images / build (push) Successful in 38s
Build and Push Docker Images / smoke (push) Successful in 0s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WpTYCBLH58XzrM7n3xPJ5N
2026-08-24 11:25:14 +00:00
claude 4e09f27e4d ci: trigger Watchtower instead of waiting for its poll
Build and Push Docker Images / build (push) Successful in 14s
Build and Push Docker Images / smoke (push) Successful in 0s
The poll is the fallback, not the mechanism. It is 30s on staging but 300s
on production, so a green run could sit five minutes ahead of the deploy it
claimed to have made -- the smoke job was absorbing that wait.

Copied from easyDent verbatim, including the soft failure: an unset secret
or an unreachable API logs and falls back to the poll rather than failing
the build. That matters right now because production has no Watchtower HTTP
API yet, so the production URL will do nothing until the infra stack there
is updated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WpTYCBLH58XzrM7n3xPJ5N
2026-08-24 11:01:35 +00:00
claude 449e4d2df3 ci: the deploy health check must follow redirects
Build and Push Docker Images / build (push) Successful in 13s
Build and Push Docker Images / smoke (push) Successful in 15s
jecreativ.ro's first production deploy went red for behaving exactly as
configured: it runs in UnderConstruction mode, so `/` correctly answers 302
to the placeholder, and the check asserted a bare 200.

Now `-L` follows the redirect and the FINAL code is asserted. A redirect
means the app is up and routing, which is what this step is for; a real
failure (502, 500, refused) still reports.

The version gate itself was right throughout -- it read 1bafc14 from the
production host, including through the under-construction middleware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WpTYCBLH58XzrM7n3xPJ5N
2026-08-24 10:34:16 +00:00
claude dee6c5a205 ci: make the branch strategy enforced instead of merely intended
Build and Push Docker Images / build (push) Successful in 20s
Build and Push Docker Images / smoke (push) Successful in 30s
Same guard as the four site repos. myAi was the ONLY stack that kept its
environment variables through the 26 July migration, so it alone deployed
correctly all along -- but it shares the ${IMAGE_TAG:-staging} default that
made the others fail silently, so it gets the same treatment.

1. ${IMAGE_TAG:-staging} -> ${IMAGE_TAG:-IMAGE_TAG-NOT-SET} on all eight
   images: an unconfigured stack fails the pull rather than quietly becoming
   staging.

2. A smoke job that asks the deploy host what commit it serves, via a
   /version.json stamped into the web image at build time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WpTYCBLH58XzrM7n3xPJ5N
2026-08-24 10:21:18 +00:00
3 changed files with 118 additions and 9 deletions
+93 -1
View File
@@ -23,6 +23,7 @@ env:
CV_SEARCH_JOB_IMAGE: apps/myai-cv-search-job
PAGE_FETCHER_API_IMAGE: apps/myai-page-fetcher-api
IMAGE_TAG: ${{ github.ref_name }} # branch name == image tag (staging | production)
WEB_PORT: "5140" # host port the web container is published on
jobs:
build:
@@ -60,7 +61,8 @@ jobs:
- name: Build Web image
run: |
docker build -f web/Dockerfile -t "${REGISTRY_HOST}/${WEB_IMAGE}:${IMAGE_TAG}" .
docker build --build-arg GIT_SHA="${{ github.sha }}" \
-f web/Dockerfile -t "${REGISTRY_HOST}/${WEB_IMAGE}:${IMAGE_TAG}" .
- name: Build CV cleanup job image
run: |
@@ -106,7 +108,97 @@ jobs:
run: |
docker push "${REGISTRY_HOST}/${PAGE_FETCHER_API_IMAGE}:${IMAGE_TAG}"
# Watchtower's poll is the fallback, not the mechanism: 30s on staging but 300s on
# production, so without this a green run can sit five minutes ahead of the deploy it
# claims to have made. Copied from easyDent, including the soft failure -- an unset
# secret or an unreachable API must degrade to the poll, never fail the build.
- name: Trigger Watchtower redeploy
env:
URL_STAGING: ${{ secrets.WATCHTOWER_URL_STAGING }}
URL_PRODUCTION: ${{ secrets.WATCHTOWER_URL_PRODUCTION }}
TOKEN: ${{ secrets.WATCHTOWER_TOKEN }}
run: |
if [ "${IMAGE_TAG}" = "production" ]; then URL="${URL_PRODUCTION}"; else URL="${URL_STAGING}"; fi
if [ -n "${URL}" ] && [ -n "${TOKEN}" ]; then
echo "Triggering Watchtower at ${URL}"
curl -sf -m 30 -H "Authorization: Bearer ${TOKEN}" "${URL}" && echo " -> redeploy triggered" \
|| echo " -> trigger failed; Watchtower will still pick it up on the next poll"
else
echo "Watchtower push-trigger not configured (WATCHTOWER_* secrets unset); relying on the poll interval."
fi
- name: Reclaim disk space (keep recent build cache)
if: always()
run: |
docker image prune -f # dangling only (keep base images)
# Building and pushing an image proves nothing about what the host is running.
# Watchtower pulls asynchronously, and for a month it was pulling a tag nobody
# intended -- with every run green, because no step ever asked the deployed site
# what it was serving. This job asks.
#
# It polls the deploy host directly on the LAN rather than the public hostname:
# the runner sits inside the network, only easysoft.ro has a staging equivalent in
# public DNS, and going direct also takes Caddy and any CDN out of the answer.
smoke:
runs-on: host
needs: build
steps:
- name: Wait for the deploy host to serve this commit
run: |
case "${{ github.ref_name }}" in
staging) HOST=192.168.1.111 ;;
production) HOST=192.168.1.101 ;;
*) echo "::error::No deploy host mapped for '${{ github.ref_name }}'."; exit 1 ;;
esac
URL="http://${HOST}:${WEB_PORT}/version.json"
echo "Polling ${URL} for ${{ github.sha }}"
# 10 minutes: Watchtower's poke is fire-and-forget with a 30s fallback poll,
# and the container still has to start.
# ⚠️ Steps run under `bash -e -o pipefail`, so a polling loop has to be written
# defensively: the FIRST miss is the normal case, not an error.
# - `curl -sf | sed` fails the whole pipeline under pipefail while the old
# container is still up (404/connection refused), so `|| GOT=""` is required
# - `[ test ] && { ... }` returns non-zero when the test fails, which under -e
# aborts the step. Use `if`.
# Getting both wrong made the first run of this job fail in 20 seconds.
DEADLINE=$(( $(date +%s) + 600 ))
while :; do
GOT=$(curl -sf -m 15 "${URL}" 2>/dev/null | sed -n 's/.*"version":"\([^"]*\)".*/\1/p') || GOT=""
if [ "${GOT}" = "${{ github.sha }}" ]; then
echo "Serving ${GOT}."
break
fi
if [ "$(date +%s)" -ge "${DEADLINE}" ]; then
echo "::error::Timed out after 10m. ${HOST} is serving '${GOT:-nothing}', wanted ${{ github.sha }}."
echo "Either Watchtower never pulled the new image, the container failed to"
echo "start, or the stack's IMAGE_TAG does not match this branch."
exit 1
fi
echo " still serving '${GOT:-nothing}' ..."
sleep 15
done
- name: Check the site actually answers
run: |
case "${{ github.ref_name }}" in
staging) HOST=192.168.1.111 ;;
production) HOST=192.168.1.101 ;;
esac
# `-L` follows redirects and we assert on the FINAL code, because a 302 from `/`
# is a healthy answer for a site running in UnderConstruction mode -- it means the
# app is up and routing. Asserting a bare 200 failed jecreativ.ro's first
# production deploy for doing exactly what it was configured to do.
#
# `|| CODE=000` for the same reason as above: curl exiting non-zero on a
# connection failure must produce a reportable code, not kill the step before
# it can say what went wrong. (`-s` without `-f` already tolerates 4xx/5xx.)
CODE=$(curl -sL -o /dev/null -w '%{http_code}' -m 20 "http://${HOST}:${WEB_PORT}/") || CODE=000
if [ "${CODE}" != "200" ]; then
echo "::error::Home page returned ${CODE}."
exit 1
fi
echo "Home page 200."
+16 -8
View File
@@ -1,6 +1,14 @@
# ⚠️ The IMAGE_TAG fallback is a DELIBERATELY INVALID tag, not `staging`.
# On 2026-07-26 these stacks were recreated by hand and lost their environment
# variables. The old `${IMAGE_TAG:-staging}` then quietly resolved to `staging`, so the
# production host pulled staging images -- with no mail credentials and no recipient
# addresses -- and served them for a month. Nothing failed, because falling back to a
# real tag is indistinguishable from being configured. Now an unset IMAGE_TAG yields
# `IMAGE_TAG-NOT-SET`, the pull fails with "manifest not found", the running container
# is left untouched and the deploy goes red. Loud beats plausible.
services:
rag-api:
image: registry.easysoft.ro/apps/myai-rag-api:${IMAGE_TAG:-staging}
image: registry.easysoft.ro/apps/myai-rag-api:${IMAGE_TAG:-IMAGE_TAG-NOT-SET}
container_name: myai-rag-api
environment:
- ASPNETCORE_ENVIRONMENT=${ASPNETCORE_ENVIRONMENT:-Staging}
@@ -50,7 +58,7 @@ services:
- "com.centurylinklabs.watchtower.enable=true"
cv-matcher-api:
image: registry.easysoft.ro/apps/myai-cv-matcher-api:${IMAGE_TAG:-staging}
image: registry.easysoft.ro/apps/myai-cv-matcher-api:${IMAGE_TAG:-IMAGE_TAG-NOT-SET}
container_name: myai-cv-matcher-api
depends_on:
- rag-api
@@ -102,7 +110,7 @@ services:
- "com.centurylinklabs.watchtower.enable=true"
email-api:
image: registry.easysoft.ro/apps/myai-email-api:${IMAGE_TAG:-staging}
image: registry.easysoft.ro/apps/myai-email-api:${IMAGE_TAG:-IMAGE_TAG-NOT-SET}
container_name: myai-email-api
environment:
- ASPNETCORE_ENVIRONMENT=${ASPNETCORE_ENVIRONMENT:-Staging}
@@ -143,7 +151,7 @@ services:
- "com.centurylinklabs.watchtower.enable=true"
api:
image: registry.easysoft.ro/apps/myai-api:${IMAGE_TAG:-staging}
image: registry.easysoft.ro/apps/myai-api:${IMAGE_TAG:-IMAGE_TAG-NOT-SET}
container_name: myai-api
depends_on:
- cv-matcher-api
@@ -217,7 +225,7 @@ services:
- "com.centurylinklabs.watchtower.enable=true"
cv-cleanup-job:
image: registry.easysoft.ro/apps/myai-cv-cleanup-job:${IMAGE_TAG:-staging}
image: registry.easysoft.ro/apps/myai-cv-cleanup-job:${IMAGE_TAG:-IMAGE_TAG-NOT-SET}
container_name: myai-cv-cleanup-job
depends_on:
- api
@@ -247,7 +255,7 @@ services:
- "com.centurylinklabs.watchtower.enable=true"
cv-search-job:
image: registry.easysoft.ro/apps/myai-cv-search-job:${IMAGE_TAG:-staging}
image: registry.easysoft.ro/apps/myai-cv-search-job:${IMAGE_TAG:-IMAGE_TAG-NOT-SET}
container_name: myai-cv-search-job
depends_on:
- cv-matcher-api
@@ -300,7 +308,7 @@ services:
- "com.centurylinklabs.watchtower.enable=true"
page-fetcher-api:
image: registry.easysoft.ro/apps/myai-page-fetcher-api:${IMAGE_TAG:-staging}
image: registry.easysoft.ro/apps/myai-page-fetcher-api:${IMAGE_TAG:-IMAGE_TAG-NOT-SET}
container_name: myai-page-fetcher-api
environment:
- ASPNETCORE_ENVIRONMENT=${ASPNETCORE_ENVIRONMENT:-Staging}
@@ -332,7 +340,7 @@ services:
- "com.centurylinklabs.watchtower.enable=true"
web:
image: registry.easysoft.ro/apps/myai-web:${IMAGE_TAG:-staging}
image: registry.easysoft.ro/apps/myai-web:${IMAGE_TAG:-IMAGE_TAG-NOT-SET}
container_name: myai-web
depends_on:
- api
+9
View File
@@ -16,4 +16,13 @@ EXPOSE 8080
ENV ASPNETCORE_URLS=http://0.0.0.0:8080
COPY --from=build /app/publish .
# Stamp the commit into the image so a deploy can be verified from the outside.
#
# Without this a smoke test can only ask "does the site return 200?" -- which it did
# throughout the month production was quietly serving staging images. A status check
# cannot tell one build from another; /version.json can, so the smoke job refuses to
# pass until the host is actually serving THIS commit.
ARG GIT_SHA=unknown
RUN mkdir -p wwwroot && printf '{"version":"%s"}' "$GIT_SHA" > wwwroot/version.json
ENTRYPOINT ["dotnet", "web.dll"]