Terraform Associate: data sources — depends_on and read-only edges
Managed resources write. Data sources read. depends_on is the rare explicit edge.
<!-- hal:authoritative:yaml -->
*Managed resources write. Data sources read. depends_on is the rare explicit edge when the graph cannot see the relationship.*
§I - Frame: Associate objectives for tonight
09-20 covered dynamic blocks, for_each, count, and splat. Leave expansion grammar.
09-17 covered ephemeral vs sensitive vs write-only. Leave secrets.
09-14 covered import and moved. Leave brownfield addressing.
09-11 covered remote backends and partial config. Leave state location.
09-08 covered check blocks and testing. Leave assertions.
09-05 covered workspaces. Leave workspace selection.
Tonight the Associate stems ask how configuration reads the world and how dependency edges form: data sources, the read-only edge in the graph, and when depends_on is justified. Ops already shows Cloud Run plus IAM members (managed writes). Dev inventories provider schema. Cert names the exam vocabulary for reads and edges.
§II - Four claims
Claim one. Data sources read existing objects. A data block queries a provider API (or another read path) and exports attributes for use elsewhere. It does not create or destroy the looked-up object. Lab 17 trains cross-stack lookups and removing hardcoded ids.
Claim two. Data values are often known only after refresh/plan conversation with the API. Some attributes are available at plan time; some stay unknown until apply in more complex graphs. Associate stems care that you do not treat a data source like a managed resource you can "update in place" through the data block.
**Claim three. Implicit edges beat ritual depends_on.** If resource B interpolates data.x.y.id or resource.a.id, Terraform already has the edge. Prefer interpolation. Reach for depends_on when the relationship is real but invisible in expressions (for example, a policy that must exist before an API call that does not take that policy id as an argument).
**Claim four. depends_on does not change create/destroy order magic beyond the edge it adds.** It does not replace lifecycle rules. It does not fix provider bugs. Overusing it hides missing interpolations and makes graphs harder to read.
§II.b - Read-only edges beside Ops and Dev
Ops tonight manages Cloud Run and IAM members. Those are write-side addresses. A data source might still appear in the same root module: for example, looking up a project number, a Secret Manager secret metadata id (not the secret payload), or an Artifact Registry repository name that already exists outside this stack. The Associate distinction is whether the block is resource or data, not whether the cloud product is "important."
Dev tonight reads provider schema JSON. That is CLI introspection, not a Terraform data source. Do not call terraform providers schema -json a data source on the exam. Data sources live in configuration. Schema JSON is a CLI diagnostic.
When a stem shows depends_on = [data.google_project.this], ask whether an interpolation already referenced that data source. If yes, the depends_on is redundant. If the managed resource never interpolates the data object yet still requires it to exist first for an invisible side effect, depends_on may be justified. Prefer stems that force that judgment over stems that only define vocabulary.
Lab 20 (external data shaping) stays referenced, not primary. External data is a specialized read path; Associate foundational coverage is ordinary provider data sources first.
§III - Exam traps
- Using a data source to "create" infrastructure: wrong shelf. Managed
resourcecreates;datareads. - Hardcoding ids when a data source filter exists: Lab 17 success criterion rejects leftover hardcoded infrastructure ids.
- **
depends_oneverywhere "for safety":** noise. Prefer expression edges. - **Confusing
terraform_remote_statedata with backend config:** remote state data reads another state's outputs; backend config chooses where this state lives (already filed 09-11). - **Treating IAM
allUserson Cloud Run as a data-source problem:** that is Ops managed IAM. Data sources might look up a service account email; they do not replace the invoker binding.
§IV - Drill (Lab 17 primary)
- Open Lab 17 (
the study notes/tfpro-labs/labs/17-data-sources-cross-stack-broken/). Replace hardcoded ids with data sources. Surface looked-up values in outputs. - Draw one sentence: which values are known after plan, and which still depend on remote reads.
- Find one place in a throwaway module where
depends_onwould be wrong (interpolation already exists) and one place where it would be justified (side effect with no argument edge). - Answer three stems: data vs resource; when
depends_onis warranted; why hardcoded ids fail Lab 17.
Success: Lab 17 plan path clean; you can explain a read-only edge without inventing a managed resource; you can refuse ritual depends_on.
§V - Close instruction
File three flash lines: data reads / resource writes; interpolation before depends_on; Lab 17 kills hardcoded ids. Maghrib owns quiz.html. Pair: Ops Cloud Run + run.invoker; Dev providers schema JSON census.
Related
- Ops: Cloud Run + run.invoker
- Dev: providers schema census
- Prior Cert: dynamic/for_each (09-20)
- Bootcamp Lab 17