Gartner's research points to a finding most AI budgets still ignore: the projects that fail are rarely let down by the model. They are let down by the data. Gartner predicted that through 2026, organisations will abandon 60% of AI projects that are not supported by AI-ready data.
For two years the default fix was to copy everything into the AI platform: export the CRM, sync the data warehouse, index the shared drives. A different pattern is now going mainstream. It is called zero-copy, and in October 2026 Google made it the recommended way for Gemini Enterprise to reach BigQuery. Here is what it is, how it works, and when it is the right call for a Hong Kong enterprise.
What is zero-copy AI?
Zero-copy AI is an architecture in which AI agents query enterprise data where it already lives, in your databases, warehouses and business systems, instead of copying it into a separate AI data store. The agent sends a governed query, receives only the answer it needs, and the source system stays the single version of truth.
The term comes from data engineering, where zero-copy sharing lets two platforms read the same table without duplicating it. Applied to AI, it changes the question from "how do we move our data to the AI?" to "how do we let the AI ask our data a question safely?"
Three building blocks usually sit underneath it:
--- Query federation: one request is split across several source systems, each answers its part, and the results are combined.
--- A metadata layer: a catalogue that tells the agent which tables exist, what the columns mean and who may see them.
--- A standard connector: increasingly the Model Context Protocol (MCP), which gives agents a consistent way to call a database tool.
Why is zero-copy suddenly on the enterprise agenda in 2026?
Zero-copy moved from data-team jargon to board-level AI strategy because the major platforms adopted it within months of each other. Google, Salesforce and SAP now offer ways for AI to read live business data in place, and agents need fresh, permissioned data far more than chatbots ever did.
The signals are concrete. In October 2026, Google's Gemini Enterprise documentation introduced a federated query mode, in preview, for BigQuery, Spanner, Cloud SQL and AlloyDB. It queries data in place through first-party MCP servers, using each user's own credentials, without ingesting anything into a Gemini data store. Google recommends it for BigQuery because it avoids egress costs from moving data.
Salesforce has built its agent strategy on the same idea. Its CTO describes Zero Copy as critical to a CIO's agentic AI strategy, with partners including Snowflake, Databricks, Google BigQuery and AWS Redshift. SAP and Google Cloud have announced zero-copy sharing between SAP Business Data Cloud and BigQuery.
The underlying driver is data readiness. A Gartner survey of 248 data management leaders found that 63% of organisations either do not have, or are unsure they have, the right data management practices for AI. Copying more data into more places does not close that gap. It usually widens it.
How does zero-copy data access actually work?
A zero-copy agent works in three steps: it looks up what data exists and what it means, it writes a query using that context, and it runs the query against the live source under the asking user's permissions. Only the result travels back to the AI, not the underlying dataset.
Google's published design for Gemini Enterprise is a useful reference model, because it shows each step explicitly:
--- Step 1, authenticate as the user: each connector is an MCP server that executes queries with the requester's own IAM permissions or OAuth credentials.
--- Step 2, discover with metadata: a knowledge catalogue returns relevant tables, schemas, business glossary terms and historical query patterns.
--- Step 3, query live: the agent formulates SQL and returns governed, real-time results.
Contrast this with the ingestion model most firms used for document AI, covered in our enterprise AI knowledge base comparison. Ingestion copies and indexes content in advance. That works well for policies and contracts that change monthly. It works poorly for inventory, claims status or cash positions that change every minute.
Is zero-copy safer than copying data into an AI platform?
Zero-copy is usually safer on data sprawl, because there is no second copy to secure, retain or delete. It is not automatically safer on access. If the agent inherits a user's broad database permissions, it can reach everything that user can reach, so permission hygiene becomes the real control.
The security case is straightforward. Every export to a new tool expands the attack surface, creates another retention schedule and another place a data subject access request must reach. Under Hong Kong's Personal Data (Privacy) Ordinance, Data Protection Principle 4 requires all practicable steps to protect personal data against unauthorised access. Fewer copies make that easier to evidence.
The caveat comes from Google itself. Its documentation states plainly that Data Cloud connectors are not scoped to specific resources. A connector attached to an app can access all resources in that data source that the authenticated user has permission to view. If your finance analysts have historically had read access to the entire warehouse "just in case", an agent acting as them now has it too.
This is why zero-copy and AI agent identity are the same conversation. The agent's access is only as tight as the identity it borrows.
When should you copy data instead of querying it in place?
Copy data when the content is unstructured, changes slowly, or must be curated for accuracy. Query in place when the data is structured, changes frequently, is sensitive, or is too large to duplicate economically. Most enterprises will run both, so the useful decision is which pattern each workload gets.
Four questions settle the choice for any single use case:
--- How fresh must the answer be? Minutes or hours points to zero-copy. Weeks points to ingestion.
--- Is the data structured? Tables and transactions suit live queries. PDFs, emails and slide decks still need indexing.
--- How sensitive is it? Personal data, client positions and pricing favour leaving data at source.
--- How much accuracy can you tolerate losing? Free-form queries on complex schemas are less precise than a curated dataset.
The last question matters more than vendors admit. Google notes that federated queries on relational databases have no table scoping or verified queries, so accuracy can vary on complex transactional schemas. For high-stakes metrics it recommends a separately scoped data agent with curated tables and verified SQL.
What does zero-copy look like in a Hong Kong enterprise?
In practice, zero-copy shows up as an operations or finance question answered from live systems in seconds, without a new data pipeline. The pattern fits logistics, insurance and property management well, because each runs fast-changing structured data that is costly and risky to duplicate.
Scenario 1: a regional logistics group
A 400-person freight forwarder keeps shipment events in a cloud warehouse. Customer service wants an agent that answers "where is consignment X and is it at risk of missing the vessel cut-off?" A nightly copy would be stale by morning. A zero-copy query against the live table answers correctly at 3pm, and the shipment data never leaves the warehouse.
Scenario 2: an insurer's claims operation
A claims manager asks how many motor claims above HK$200,000 have been open for more than 60 days. The data includes personal information. Querying in place under the manager's own permissions keeps the PDPO position simple: no new copy, no new retention policy, and the existing access log still applies.
Scenario 3: a property management company
Repair tickets, contractor invoices and estate budgets sit in three systems. Federation lets an agent join them for a monthly variance question without building a new data mart. The trade-off is latency: cross-system queries can take longer than a pre-built report.
What are the most common zero-copy pitfalls?
The most common pitfalls are poor metadata, over-broad permissions, uncontrolled query costs and overconfidence in answers from complex schemas. None of these is a technology failure. Each is a governance gap that copying data used to hide and that zero-copy now exposes.
--- Undocumented data: an agent cannot interpret a column called FLD_07. Without a maintained glossary, live queries return confident nonsense.
--- Inherited over-permissioning: years of generous access grants become agent capabilities overnight.
--- Runaway query bills: every agent question can trigger a warehouse scan. Set a default billing project and budget alerts before launch.
--- Source system load: analytical queries against a production database can slow the systems your staff depend on.
--- Unverified metrics: "revenue" may mean three different things across systems. Board-level numbers need verified queries, not improvisation.
A useful test is to ask whether your data team could answer the agent's question manually in under an hour. If not, the agent will not do better. It will simply be wrong faster.
How should enterprise leaders evaluate zero-copy claims from vendors?
Evaluate zero-copy claims by asking where queries run, whose identity they use, how access is scoped, what is logged and how cost is controlled. A vendor that cannot answer all five in writing is selling a connector, not a governed zero-copy architecture.
Use these five questions in your next vendor meeting or RFP:
--- Where does the query execute, and does any raw data leave the source system or Hong Kong?
--- Whose credentials does the agent use, the end user's or a shared service account?
--- Can access be scoped to named datasets and tables, or is it all-or-nothing?
--- What is logged, including the generated query, the result size and the user, and for how long?
--- How are costs capped per user, per agent and per month?
Ask for a demonstration on your own schema, not the vendor's sample data. Zero-copy demos on clean tables are easy. The real test is your 15-year-old ERP.
Conclusion: stop moving data, start governing access
Zero-copy does not make data readiness optional. It makes data readiness the whole game. The organisations that benefit first will be those that already know what their data means and who should see it. For everyone else, the first investment is a catalogue and an access review, not another pipeline.
The strategic takeaway is simple: decide, workload by workload, which data the AI should read in place and which it should copy, and make permissions, not plumbing, the centre of the design.
We understand AI. We understand you. With UD by your side, AI never feels cold.
Reviewed by the UD enterprise AI team. Platform details were checked on 8 October 2026 against Google Cloud documentation and the vendor and analyst sources linked above. Preview features change often, so confirm current status before you commit.
Find Out Whether Your Data Is Ready for AI Agents
Now that you have the framework, the next step is identifying which of your workflows have data an agent can safely use today. Start with the free AI Ready Check. We'll walk you through every step, from data readiness and access review to pilot design and performance tracking.