Memory hook: Code, data, environment: version all three.
Must remember
- An Azure Machine Learning workspace organizes jobs, models and assets. Datastores describe supported storage connections; data assets version references to data; neither automatically guarantees that underlying mutable data stays unchanged.
- Compute instances suit development; compute clusters and other supported targets run scalable jobs. Configure identity, networking, quota and auto-shutdown/scale-down to match workload and cost requirements.
- Environments define runtime dependencies; components package reusable steps; registries share supported assets across workspaces. Pin versions/digests so a pipeline can be reproduced.
- Use Bicep/Azure CLI and reviewed configuration to provision infrastructure. Separate development/test/production and pass environment-specific settings without hard-coded credentials.
- GitHub Actions can authenticate through federated identity rather than stored long-lived secrets. Grant the workflow identity only required deployment and data permissions.
- Private endpoints, managed networking and storage access must work together. A private workspace alone does not prove that compute can reach every required package, registry or datastore.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Reuse a tested transformation across teams | A versioned component/environment in a supported registry. |
| Provision identical workspaces repeatedly | Reviewed IaC with separate identities and approved environment parameters. |
Traps
- A versioned asset pointing to overwritten files is not truly reproducible data.
- Network isolation can break dependency downloads if required paths are not designed.
Active recall
1. What does an environment capture?
Runtime dependencies and container configuration used to execute code.
2. Why use federated CI identity?
To obtain scoped short-lived access without distributing a permanent secret.
3. How distinguish datastore and data asset?
A datastore describes access to storage; a data asset provides a named/versioned data reference.
4. Why separate workspaces/environments?
To control access, cost, experimentation and production change boundaries.
5. What should infrastructure validation check?
Permissions, networking, storage, compute, quotas and repeatable provisioning.