Python SQS DLQ and redrive policy census
A queue without a redrive policy fails forever in place. A DLQ without depth visibility fails silently.
<!-- hal:authoritative:yaml -->
A queue without a redrive policy fails forever in place. A DLQ without depth visibility fails silently.
§I — Frame
Python-day Ops against a fresh AWS family. CloudWatch alarm census (09-10), GCS lifecycle (09-13), and Azure VNet/NSG (09-19) stay on their shelves. EventBridge bus/rule/target (08-29) stays too.
Today the concrete surface is SQS redrive. SAP bootcamp names Dead Letter Queues as the place unprocessed messages land after maxReceiveCount retries, with ReceiveCount incremented on each receive and the original enqueue timestamp preserved when the message moves. Ops prints that wiring. It does not purge. It does not rewrite attributes. It does not start a message-move task.
§II — Three shelves on one queue family
| Shelf | What it does | What it does not do |
|---|---|---|
| Source queue | Holds work; increments ReceiveCount on each receive | Guarantee success after N receives |
| RedrivePolicy | Names deadLetterTargetArn + maxReceiveCount | Exist by default (unset means no DLQ) |
| Dead-letter queue | Receives poison after threshold; longer retention is usual | Fix the consumer bug by itself |
SAP notes: Standard is at-least-once with best-effort order; FIFO is exactly-once ordered with .fifo suffix and throughput caps. Both can attach a DLQ. A single DLQ can serve multiple sources. Visibility timeout and retention are separate knobs from redrive.
§III — Mechanism: the census, not the remediation
Reuse a regional boto3.client("sqs"). Prefer get_queue_attributes over inventing fields. Attribute names that matter today: QueueArn, RedrivePolicy, ApproximateNumberOfMessages, ApproximateNumberOfMessagesNotVisible, ApproximateNumberOfMessagesDelayed, FifoQueue, VisibilityTimeout, MessageRetentionPeriod.
import json
from dataclasses import dataclass
from typing import Iterator, Optional
import boto3
from botocore.exceptions import ClientError
@dataclass(frozen=True)
class RedriveRow:
queue_url: str
queue_arn: str
fifo: bool
max_receive: Optional[int]
dlq_arn: Optional[str]
visible: int
not_visible: int
delayed: int
dlq_visible: Optional[int]
ATTRS = [
"QueueArn",
"RedrivePolicy",
"ApproximateNumberOfMessages",
"ApproximateNumberOfMessagesNotVisible",
"ApproximateNumberOfMessagesDelayed",
"FifoQueue",
"VisibilityTimeout",
"MessageRetentionPeriod",
]
def _int_attr(attrs: dict, key: str) -> int:
return int(attrs.get(key, "0") or "0")
def census(client) -> Iterator[RedriveRow]:
paginator = client.get_paginator("list_queues")
for page in paginator.paginate():
for url in page.get("QueueUrls") or []:
try:
resp = client.get_queue_attributes(QueueUrl=url, AttributeNames=ATTRS)
except ClientError as exc:
code = exc.response.get("Error", {}).get("Code", "")
if code in {"AccessDenied", "AWS.SimpleQueueService.NonExistentQueue"}:
continue
raise
attrs = resp.get("Attributes") or {}
policy_raw = attrs.get("RedrivePolicy")
max_receive = None
dlq_arn = None
if policy_raw:
policy = json.loads(policy_raw)
max_receive = int(policy["maxReceiveCount"])
dlq_arn = policy.get("deadLetterTargetArn")
dlq_visible = None
if dlq_arn:
# Resolve DLQ URL once per row; skip if the target is gone or denied.
try:
dlq_url = client.get_queue_url(
QueueName=dlq_arn.rsplit(":", 1)[-1]
)["QueueUrl"]
dlq_attrs = client.get_queue_attributes(
QueueUrl=dlq_url,
AttributeNames=["ApproximateNumberOfMessages"],
)["Attributes"]
dlq_visible = _int_attr(dlq_attrs, "ApproximateNumberOfMessages")
except ClientError:
dlq_visible = None
yield RedriveRow(
queue_url=url,
queue_arn=attrs.get("QueueArn", ""),
fifo=attrs.get("FifoQueue", "false").lower() == "true",
max_receive=max_receive,
dlq_arn=dlq_arn,
visible=_int_attr(attrs, "ApproximateNumberOfMessages"),
not_visible=_int_attr(attrs, "ApproximateNumberOfMessagesNotVisible"),
delayed=_int_attr(attrs, "ApproximateNumberOfMessagesDelayed"),
dlq_visible=dlq_visible,
)
def print_census(rows: list[RedriveRow]) -> None:
bare = [r for r in rows if r.dlq_arn is None]
armed = [r for r in rows if r.dlq_arn is not None]
hot = [r for r in armed if (r.dlq_visible or 0) > 0]
print(f"queues={len(rows)} bare={len(bare)} armed={len(armed)} dlq_hot={len(hot)}")
for r in sorted(armed, key=lambda x: (-(x.dlq_visible or 0), x.queue_arn)):
print(
f"{r.queue_arn} fifo={r.fifo} maxReceive={r.max_receive} "
f"src_vis={r.visible} dlq={r.dlq_arn} dlq_vis={r.dlq_visible}"
)
Three rules fall out.
**Rule one. Unset RedrivePolicy is a finding, not a default.** Bare queues keep failing in place after VisibilityTimeout returns the same poison.
Rule two. Depth on the DLQ is the operator signal. SAP notes you alarm on DLQ failure. Approximate counts are eventually consistent; treat them as triage, not exact ledger.
Rule three. Retention on the DLQ should outlive the source. Bootcamp: the enqueue timestamp does not reset when a message moves. A short DLQ retention deletes evidence before you read it.
FIFO detail: source and DLQ should both be FIFO (or both Standard). Mixing types is a configuration error you catch by printing FifoQueue on both ARNs, not by guessing from the name alone.
§IV — What not to do
Do not call purge_queue. Do not set_queue_attributes to "fix" redrive in the census tool. Do not start_message_move_task from this script (that is remediation, Maghrib-adjacent). Do not treat ApproximateNumberOfMessages as Exact. Do not skip FIFO suffix checks when the ARN ends in .fifo. Do not re-teach EventBridge schedules (08-29) or CloudWatch alarm shapes (09-10) as the primary drill.
§V — Close instruction
Run the census against one account/region. Confirm every production source either has RedrivePolicy or is explicitly listed bare. Pick one armed queue with dlq_visible > 0 and open the DLQ messages by hand (Console or a separate receive script). Pair: Dev teaches singledispatch for rendering census rows by type; Cert opens Direct Connect resiliency (not PrivateLink, not this queue).