Skip to main content

Blog system: content, publishing, and scheduling

Refract’s blog stores articles in Postgres, serves MDX from the CDN, and exposes public GraphQL to the marketing site. Scheduled publishes use a reconciliation controller plus a delayed queue — the database is always the source of truth.

How the blog is built

Editor markdown syntax lives on Blog article markdown cheatsheet.

Article and media CDN lifecycle

Article MDX and blog media are public, cacheable objects. We version by object key — not in-place overwrite — so browsers and edge CDNs can cache aggressively without serving stale content after an edit. Why two ops per article save: overwriting the same URL/key would be one PutObject, but anything already cached at the old URL could keep serving old MDX until TTL expiry or an explicit purge. A new key + DB pointer update gives readers a fresh URL; deleting the previous key limits orphan storage. Billing implication: on Cloudflare R2 (and similar object stores), PutObject and DeleteObject are billed as mutating request classes — not storage size alone. Heavy editorial churn (many Saves on long drafts) increases request cost linearly with saves, even when byte size is small. Explicit Save in the CMS reduces accidental churn from autosave; it does not change the per-save cost when you do save. Orphan objects can remain under articles/ or blog-media/ if a save failed mid-flight, a row was deleted without the cleanup path, or code predated delete-on-replace. The database (content_path, blog_media.cdn_object_key) is the source of truth for what is live. See CDN overview — object lifecycle and Cloudflare R2 — blog operations.

Scheduling and publishing

These contracts are intentional — not missing functionality.
Scheduling is intentionally bounded. The queue is a rolling execution buffer; the database holds long-term schedule state.
DLQ is not part of the recovery path. Failed publish jobs are observable in the DLQ for debugging only. Recovery is via reconciliation controllers and the CMS reschedule hook.
  • Queue: BLOG_ARTICLE_PUBLISH, jobId = blog-publish-{articleId} (no : — BullMQ constraint), replace: true on enqueue
  • Cancel: removeQueuedJob — skips active jobs (skipped_active_cancel, no throw)
  • Consumer → publishBlogArticle atomic UPDATE (status, scheduled_at, archived_at)
  • Discovery regen on every visibility change; 6h blogRegenerateDiscovery cron is backstop only

Monitoring and alerting

Reconcile and publish paths emit low-cardinality metrics from apps/backend/src/utilities/blog/metrics.ts. Wire dashboards or alerts on these counters — not on per-article IDs. Runbook (stuck scheduled article):
  1. Confirm row in blog_articles: status = scheduled, scheduled_at in the past, archived_at null.
  2. Check consumer logs for blogArticlePublish and metric blog.publish.consumer.
  3. Trigger reconcile: wait for daily blogPublishScheduled / weekly blogPublishReconcile, or reschedule from the CMS publish drawer (re-enqueues via publishEnqueue).
  4. If enqueue keeps failing, inspect queue DLQ for payload debugging only — recovery is reconcile + CMS reschedule, not DLQ replay.

Operator mental model

  • Debug DB first (status, scheduled_at, archived_at)
  • The queue is not schedule storage
  • Watch for stuck scheduled rows past scheduled_at, growing queue depth, or repeated blogRegenerateDiscovery warn logs

Extending this system

  1. Add business logic in apps/backend/src/utilities/blog/
  2. Register queue name + DLQ in apps/backend/src/utilities/queue.ts and config (development.ts, staging.ts, production.ts)
  3. Add consumer in apps/backend/src/tools/queue/consumers/ and register in consumerRegistry.ts
  4. Add schedule name in scheduleName.ts, handler in schedules/, and register in scheduleRegistry.ts
  5. Write tests under apps/backend/src/utilities/blog/__tests__/
  6. Update this page when behavior changes

What not to do

❌ Never publish inline in scheduler handlers — call reconcileBlogPublishJobs only.
❌ Never use the queue as the source of truth for schedules.
❌ Never replay DLQ messages for normal recovery.
❌ Never call sendToQueue from GraphQL — use enqueueBlogArticlePublish / cancelBlogArticlePublishJob.

File conventions

What’s next?