The state is the first problem
Terraform records which real resources belong to which code in its state. During initial local experiments, that file often lives on the developer’s computer. In a team, every run needs access to the correct state; uncoordinated simultaneous writes can corrupt it.
The answer is remote state: the state lives in a shared backend, such as Azure Blob Storage or S3. Choose a backend with supported and enabled state locking. S3 backend locking must be enabled explicitly. Versioning also helps recover accidentally deleted or damaged state. These principles apply to Terraform and OpenTofu; the settings depend on the backend and version in use.
State can contain sensitive resource attributes, such as database passwords or API tokens. Backend access, encryption and logging therefore belong in the security design. State files and plan artifacts must not be committed to Git or exposed in publicly readable pipeline output.
One state per environment, not one for everything
A single state for the whole organisation couples many changes and increases the potential impact of a mistake. Access can also expose sensitive production data. Permissions on state and permissions to change cloud resources must be secured separately.
A useful separation follows environments and ownership. Development, test and production receive separate states with appropriate permissions. Routine write access to production state and resources belongs to the production pipeline; emergency access is governed separately. Further states for areas such as networking, clusters and applications keep changes manageable. Publish shared values explicitly where possible: the terraform_remote_state data source requires access to the entire state snapshot, even when only outputs are read.
Terraform CLI workspaces also separate states, but within the same configuration and the same backend. For environments with different access rights, separation through dedicated backends is the more robust choice.
Modules: Git sources or a registry
Modules are the unit in which repetition pays off: a reviewed pattern for a VNet, a cluster, a database that all teams use, instead of each inventing it once themselves.
There are two common paths for distribution. Git sources are the
simple entry point: the module lives in its own repository, the
consumers reference it through the source and pin a tag with the
ref parameter. Exactly this pinning matters: without
ref, Terraform pulls the default branch, and that
branch moves.
An internal registry pays off as soon as several teams use the same modules. It brings version constraints instead of fixed tags, a searchable overview and documentation attached to the module. The price is another piece of infrastructure someone has to operate. For a single platform team, tagged Git sources are usually enough; for an organisation with many consumers, the balance tips.
In both cases: modules are versioned like software. Every change is tagged, consumers update deliberately, and a change that breaks existing calls gets a new major version and a note in the changelog.
Review: the plan belongs in the merge request
A code review alone is not enough with Terraform. The diff shows
the intent, the plan shows the effect, and the two can lie far
apart: an inconspicuous change to one argument can trigger a
replace that recreates a database.
The pipeline generates a plan for each merge request. Review covers both code and planned resource changes. After the merge, the pipeline applies the approved saved plan. If changes to the code or environment require a new plan, that plan is reviewed again. Plan artifacts remain access-controlled because they can contain sensitive values.
To keep reviews manageable, changes should be small: one merge request per concern. The more resources a plan affects, the harder it becomes to assess interactions. This is also a reason to separate states by ownership and lifecycle.
Policy checks also build on the plan: rules such as allowed regions, mandatory tags or forbidden public IP addresses can be checked with tools like Open Policy Agent against the machine readable plan output before a person starts the review. What the machine can reject, the reviewer does not have to hunt for. That is the same logic as Azure Policy in the landing zone, just one step earlier in the process.
Drift: when reality diverges
Drift occurs when someone changes resources outside the code, for example quickly in the portal during an incident. The state then matches neither reality nor the code, and the next apply can revert changes or replace resources nobody expected.
A terraform plan -refresh-only shows changes to managed resources and outputs without applying them. A scheduled run with notifications can expose drift. It does not comprehensively cover unmanaged resources or properties the provider does not track. For an intentional change, reconcile code and state deliberately; otherwise, review a normal plan for the correction before applying it.
The most effective lever against drift, however, is organisational: when write access in production sits with the pipeline and not with people, there are simply fewer ways to work around Terraform. Emergency access remains possible, but it is then the documented exception instead of everyday practice.
Terraform in a team is less a tooling topic than a process topic: a remote state with locking, separated per environment and building block, modules with pinned versions and a plan that someone reads before the pipeline applies it. Whoever settles these three things early has an infrastructure that is still traceable after the next team change.
What you can decide afterwards
- Where state lives, how it is locked and backed up, and how routine and emergency access are governed.
- Whether tagged Git sources are enough for your modules or whether an internal registry pays off.
- What a merge request has to look like for you so that plan, review and apply stay traceable.
Frequently asked questions
Are Terraform CLI workspaces enough to separate our environments?
Not for environments with different access rights. CLI workspaces separate states within the same configuration and the same backend. Development and production belong in dedicated backends with their own rights.
Does all of this also apply to OpenTofu?
The principles also apply to OpenTofu: shared state, active locking, versioned modules and an approved plan. Check backend settings, provider compatibility and available features for the versions in use.
Are developers no longer allowed to work locally at all?
Developers can plan locally against the development environment with appropriate permissions. Routine production changes run through the pipeline with an approved plan. Emergency access remains possible, but must be documented and reconciled with the code afterwards.
Sources
- Remote State HashiCorp, Terraform documentation
- Use modules in your configuration (module sources and versions) HashiCorp, Terraform documentation
- Manage resource drift HashiCorp, Terraform tutorial
- Remote State (OpenTofu) OpenTofu documentation
- S3 backend: locking and versioning HashiCorp, Terraform documentation
- Plan and saved plans HashiCorp, Terraform documentation
- Protect sensitive data HashiCorp, Terraform documentation
- Sharing data between configurations HashiCorp, Terraform documentation