Cloud Data Governance: Ownership, Classification, Transfers, and Retention in 2026
Practical cloud data governance: data owners, classification and discovery, GDPR transfer rules, EU Data Act switching, AI data leaks, and retention costs.

Cloud data governance answers a few plain questions: what data do we have, where is it, who owns it, who can see it, where can it go, and when do we delete it. Most organizations cannot answer all of those for every dataset, and the gaps are where breaches, fines, and runaway storage bills come from. IBM's 2026 Cost of a Data Breach report put the global average breach cost at $4.99 million, a 12% rise on the year before.
This guide covers each question with the tools and rules that apply in 2026. For privacy law obligations in more depth, see our AI and SaaS data privacy guide; for AI-specific data rules, our generative AI data governance guide.
Ownership first
Every important dataset needs a named owner in the business, accountable for who gets access and how long data is kept, and usually a steward who handles quality and documentation day to day. Platform and security teams run the controls; they should not be deciding on their own who may see customer data.
A small data council with security, legal, IT, and the main business domains sets policy and settles disputes. Keep it small enough to make decisions.
Know what you have
You cannot protect or delete data you have not found. Two kinds of tools help:
- Catalogs inventory datasets, owners, descriptions, and lineage. Examples include Microsoft Purview, AWS Glue Data Catalog and Amazon DataZone, and Google's Knowledge Catalog. Google's older Data Catalog service was set to be discontinued on June 1, 2026, so teams still on it need to finish migrating.
- Discovery and data security posture management (DSPM) tools scan storage for sensitive content, such as personal data, payment card numbers, and credentials, and flag where it sits in the wrong place. Amazon Macie does this for S3; third-party DSPM tools cover multiple clouds and SaaS.
Pay particular attention to copies: staging buckets from migrations, database snapshots, analytics extracts, test environments seeded with production data, and exports sent to vendors. These tend to be forgotten and are often less protected than the original.
Classify and apply controls
A four-level scheme works for most organizations:
| Level | Examples | Typical controls |
|---|---|---|
| Public | Published marketing, public filings | None beyond integrity |
| Internal | Policies, most business documents | Employee access, standard encryption |
| Confidential | Financial results before release, contracts, source code | Role-based access, logging, no public sharing |
| Restricted | Customer personal data, health data, payment data, credentials | Least privilege, strong authentication, key management, access reviews, DLP |
Automate where possible: tag datasets by level, then let policy enforce the controls. Keep the scheme simple enough that engineers can apply it without a lookup table.
Access
- Grant access through roles or attributes, not individuals.
- Review access to restricted data at least quarterly and remove what is not used.
- Prefer short-lived credentials over long-lived keys.
- Keep S3 Block Public Access on. AWS has enabled it by default for new buckets since April 2023, and account-level settings stop exceptions from slipping in.
Where data can go
Transfers under GDPR. GDPR does not require EU storage, but transfers of personal data outside the European Economic Area (Articles 44 to 49) need a legal basis: an adequacy decision, standard contractual clauses with a transfer impact assessment, or binding corporate rules. For US recipients, the EU-US Data Privacy Framework covers certified companies; the EU General Court upheld it on 3 September 2025. Record which systems send personal data abroad and on what basis.
Localization rules do exist in some sectors and countries, such as certain government, health, and financial data. Check before choosing regions, and use cloud policy to restrict storage to approved regions.
US state laws. A growing number of states have general-purpose consumer privacy laws; the IAPP keeps a tracker. Under California's regulations approved in September 2025, businesses subject to the risk assessment rules had to begin complying on January 1, 2026, and must submit an attestation and summary to the CPPA by April 1, 2028.
Vendors and AI tools. Data sent to SaaS vendors and AI services is still your responsibility. Verizon's 2026 Data Breach Investigations Report found third parties involved in 48% of breaches, and 45% of employees regular AI users on corporate devices, with 67% of AI users on those devices using non-corporate accounts. Approve AI tools with enterprise terms that exclude your data from training, and extend data loss prevention rules to AI prompts and uploads.

Leaving a provider: the EU Data Act
The EU Data Act's cloud switching rules have applied since September 12, 2025. Providers must help customers move to another provider or back on premises. Until January 12, 2027 they may still charge for the costs of switching and egress; from that date, switching charges, including egress, are removed. Governance teams should keep an exit plan for important workloads: data export formats, how long an export would take, and which services have no equivalent elsewhere.
Retention and lifecycle
Keeping data forever costs money and raises the impact of a breach or lawsuit. Build a retention schedule by record type, citing the legal or business requirement for each, then enforce it with lifecycle rules.
An illustration of the cost side, using AWS list prices for US East (N. Virginia) from the price list published September 28, 2026: 500 TB (500,000 GB) of old logs in S3 Standard at $0.023 per GB-month for the first 50 TB and $0.022 after that costs about $11,050 a month. In S3 Glacier Deep Archive at $0.00099 per GB-month, the same data costs about $495 a month. The trade-offs are retrieval times of up to 12 hours (standard) or 48 hours (bulk), retrieval fees, and a 180-day minimum storage duration, so move only data you rarely need. Deleting data you have no reason to keep costs nothing at all.

Governance as code
Manual reviews do not keep up with cloud change. Put the rules into the deployment pipeline:
- Templates (Terraform, CloudFormation, Bicep) that create storage with encryption, logging, tags, and private access by default.
- Policy checks (Open Policy Agent, Checkov, cloud-native policy services) that fail a deployment when a bucket would be public or untagged.
- Organization-level guardrails that restrict regions and block risky configurations; see our cloud cost governance guide for how these work in each cloud.
- Continuous posture monitoring to catch drift after deployment; see our CSPM guide.
A 90-day starting plan
- Name owners for your 20 most important datasets.
- Run discovery across cloud storage and major SaaS apps, and list where restricted data lives.
- Adopt the four-level classification and tag the restricted datasets first.
- Review access to restricted data and remove unused grants.
- Record your cross-border transfers and their legal basis.
- Write a retention schedule for the largest data stores and turn on lifecycle rules.
- Approve enterprise AI tools and extend DLP to AI use.
For the security architecture around these controls, see our zero trust guide.
This guide is for informational purposes only and does not constitute legal advice. Data protection obligations vary by jurisdiction and sector; consult qualified privacy counsel and security professionals.



