<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Harsh Trivedi | AI & Data Engineering]]></title><description><![CDATA[Harsh Trivedi | AI & Data Engineering]]></description><link>https://harshtrivedii.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6a39322015ed6345c6f1c5cc/87b78007-fc17-4030-9ad0-2790ee291ca0.png</url><title>Harsh Trivedi | AI &amp; Data Engineering</title><link>https://harshtrivedii.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 09 Sep 2026 14:12:17 GMT</lastBuildDate><atom:link href="https://harshtrivedii.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[The AI Adoption Gap: What Boardrooms Are Buying vs. What Employees Actually Need]]></title><description><![CDATA[There’s a quiet math problem sitting underneath the AI boom that most leadership decks don’t show: probabilistic systems are being deployed to replace deterministic work, and nobody at the top seems t]]></description><link>https://harshtrivedii.hashnode.dev/the-ai-adoption-gap-what-boardrooms-are-buying-vs-what-employees-actually-need</link><guid isPermaLink="true">https://harshtrivedii.hashnode.dev/the-ai-adoption-gap-what-boardrooms-are-buying-vs-what-employees-actually-need</guid><category><![CDATA[AI Adoption]]></category><category><![CDATA[ai-agent]]></category><category><![CDATA[agentic AI]]></category><category><![CDATA[Enterprise AI]]></category><category><![CDATA[enterprise software]]></category><dc:creator><![CDATA[Harsh Trivedi]]></dc:creator><pubDate>Sun, 02 Aug 2026 14:10:29 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/c8083593-7e21-4a0e-a36d-9bef806a932b.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There’s a quiet math problem sitting underneath the AI boom that most leadership decks don’t show: probabilistic systems are being deployed to replace deterministic work, and nobody at the top seems to have clocked the difference.</p>
<p>This isn’t an anti-AI piece. It’s a reality check.</p>
<hr />
<h3><strong>Probabilistic vs. Deterministic — the distinction nobody in the C-suite asks about</strong></h3>
<p>A large language model is, by design, a probability engine. Ask it the same question twice and you can get two different answers. That’s not a bug to be patched away — it’s the nature of the thing. Traditional software, by contrast, is deterministic: same input, same output, every time, forever.</p>
<p>Most of the actual work happening inside companies — payroll calculations, leave approvals, financial reconciliations, standard reports — is deterministic. It has a correct answer, a fixed rule set, and zero tolerance for “approximately right.” This is exactly the kind of work traditional software has solved cheaply and reliably for decades.</p>
<p>Yet a huge share of current AI initiatives are aimed squarely at this deterministic work — not because it’s a good technical fit, but because nobody wants to be the department that isn’t “doing AI.”</p>
<p>The result is systems that are probabilistic being asked to do a deterministic job, wrapped in a disclaimer that says <em>“AI can make mistakes.”</em> That disclaimer isn’t boilerplate. It’s an admission that should stop a lot of these projects at the pitch stage. Instead it gets buried in a footer, and the project ships anyway.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/bdad632c-928d-4055-b3de-2eba678121fd.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h3><strong>The shift nobody validated</strong></h3>
<p>A pattern playing out across large service organizations right now: internal data and analytics teams pivoting from dashboards to “agents” — conversational interfaces that answer questions in natural language instead of showing a chart.</p>
<p>The uncomfortable part isn’t the technology. It’s who’s driving the decision. It’s rarely the HR team, the finance team, the ops or sales function — the actual consumers of these dashboards — asking for a conversational agent. It’s the teams building them, motivated by a very reasonable but very different incentive: something demonstrable for a board update, an internal AI showcase, or a slide that proves the org is “in the race.”</p>
<p>Nobody asked the person who actually uses the dashboard every day: <em>do you want to type a question instead of glancing at a chart you’ve read a hundred times?</em> In most cases, the honest answer would be no. But that question rarely gets asked, because the initiative isn’t really being built for the end user. It’s being built for the room upstairs.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/d724d034-5297-4a12-91e3-fbe2bb57a171.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h3><strong>The irony at the center of it all</strong></h3>
<p>Here’s the part that should worry finance departments more than anyone: after all this build-out, the actual source of truth for most teams is still an Excel sheet. Not because Excel is glamorous, but because it’s free, transparent, and trusted. People can see the formula. They can audit it. There’s no black box and no per-query cost.</p>
<p>Now watch what happens: the same question that used to be answered by opening a spreadsheet — free, instant, deterministic — gets routed through an AI agent instead. The agent burns tokens, calls a model, generates a natural-language summary, and then, often, regenerates the same spreadsheet as its output. The work that used to cost nothing now has a metered price tag attached to it, incurred every single time someone asks. And this was invisible while heavy subsidies and discounted token pricing made it look free. As those subsidies wind down, the actual unit economics are starting to show — and they’re not flattering.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/5dd33713-a65c-4a34-b678-5874807d4074.png" alt="" style="display:block;margin:0 auto" />

<p>None of this is being tracked at the level it should be. Conversations with people close to large-scale AI rollouts reveal environments where thousands of agents have been spun up over time — proof-of-concepts, pilots, one-off demos — with no consolidated inventory of what’s still running, what it’s costing, or whether anyone is still using it. An agent built for a demo six months ago can still be quietly consuming budget today, and the bill only becomes visible long after it should have been shut down. That’s not an AI strategy. That’s uncontrolled sprawl with a monthly invoice.</p>
<hr />
<h3><strong>Do decision-makers even look?</strong></h3>
<p>There’s a harder question buried in all of this: when the agent produces its polished natural-language summary, does the executive actually read it and think it through — or do they just want a junior to hand them a filled Excel sheet, the same as before, so they can move on with their day?</p>
<p>If it’s the latter, the AI layer isn’t informing the decision. It’s theater sitting on top of a decision-making process that never changed. You cannot responsibly base a business call on an output that ships with a disclaimer saying it might be wrong — not when the underlying deterministic data was already sitting right there, for free, before the AI system re-derived it at a cost.</p>
<hr />
<h3><strong>The people cost nobody puts in a deck</strong></h3>
<p>There’s also a quieter toll this creates on the people actually building these systems. Someone with a data engineering background — hired to build pipelines and models — ends up asked to produce a demo video to show off in a board meeting or an internal AI event. Not because it’s their job, but because the org needs <em>something</em> to point to, and the people who can actually build the thing are also the ones roped into selling it. That work often goes unrecognized and uncompensated, sitting entirely outside the job description it was layered onto.</p>
<p>This is a symptom of the same root cause: the initiative exists to be shown, not to be used. When the primary audience for a product is a board slide instead of an end user, the people doing real technical work get quietly repurposed into marketing, and nobody adjusts their role, their scope, or their pay to reflect it.</p>
<hr />
<h3><strong>The simplest test of all: the button</strong></h3>
<p>Strip away the architecture diagrams and ROI slides and there’s a test simple enough that almost nobody applies it: if the UI already has a button for the task, why would anyone type a sentence to an AI to do it instead?</p>
<p>Applying for leave, raising a support ticket, filling a timesheet — these are one-click actions today. Replacing a click with a typed request to an agent isn’t a productivity gain; it’s added friction wearing an AI badge. End users know this instinctively, which is why adoption of these “agentic” replacements for already-solved workflows tends to be quiet and thin, no matter how good the demo looked.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/9fb07847-a521-491d-aac9-caa0f3c9b283.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h3><strong>Where this doesn’t apply</strong></h3>
<p>To be fair to the technology itself: none of this means AI has no place in the enterprise. Probabilistic systems are genuinely well-suited to genuinely probabilistic problems — unstructured document synthesis, first-draft generation, anomaly detection in noisy data, coding assistance, customer sentiment analysis, and other work where “highly likely correct with human review” is an acceptable and even valuable standard. The failure isn’t AI. The failure is applying a probabilistic tool to deterministic work purely because leadership wants a story to tell, without asking the people who’d actually use it whether they wanted it in the first place.</p>
<hr />
<h3><strong>Soaring Costs and AI Agent Sprawl</strong></h3>
<p>Meanwhile, deploying AI can be <strong>expensive</strong> and hard to track. After introductory credits and subsidies end, every AI query incurs token charges — and usage can skyrocket unexpectedly. As the <em>Wall Street Journal</em> reports, companies are already “pumping the brakes” as the <em>“soaring cost”</em> of AI compute hits the budget. Unchecked, AI agents can quietly burn millions: for example, one healthcare organization ran up over <strong>$6 million</strong> in API costs (consuming a trillion tokens) <em>before</em> the finance team even noticed what was happening. A recent survey found 85% of companies routinely miss their AI cost forecasts by more than 10%. Put bluntly, work that was once done for free in a spreadsheet now incurs a token cost, often with no extra benefit — yet companies still feel pressure to keep a fleet of AI agents running. Without strict governance, the AI rollout risks becoming a <strong>budget sink</strong> rather than a productivity gain.</p>
<hr />
<h3><strong>The reckoning</strong></h3>
<p>Token discounts and subsidized pricing masked the real cost of this approach for a while. As that support fades, the unit economics of “AI for everything” are becoming visible, and in a lot of cases the ROI simply isn’t there — because the ROI was never the point. The point was optics: a board deck, a submission, an event demo, proof that the organization is “in the race.”</p>
<p>Eventually the money question comes due. And when it does, the organizations that will have something real to show are the ones that asked a simple question before building anything: <em>does the person who has to use this every day actually want it?</em> Everyone else will be left explaining, in a board meeting, why millions went into a race that never had a finish line — and why the spreadsheet is still, quietly, doing all the real work.</p>
]]></content:encoded></item><item><title><![CDATA[Enterprise-Scale Azure Key Vault Management: Deleting, Purging, and Automating Secrets]]></title><description><![CDATA[Cleaning Up Azure Key Vault: A Business and Developer’s Guide to Deleting, Bulk-Deleting, and Permanently Purging Secrets
Part of the series AI Cloud and Data Engineering Articles. Every pipeline in t]]></description><link>https://harshtrivedii.hashnode.dev/enterprise-scale-azure-key-vault-management-deleting-purging-and-automating-secrets</link><guid isPermaLink="true">https://harshtrivedii.hashnode.dev/enterprise-scale-azure-key-vault-management-deleting-purging-and-automating-secrets</guid><category><![CDATA[Azure Key Vault]]></category><category><![CDATA[Azure Key Vault Secrets]]></category><category><![CDATA[keyvault]]></category><category><![CDATA[Key Vault Secrets Management]]></category><category><![CDATA[#microsoft-azure]]></category><dc:creator><![CDATA[Harsh Trivedi]]></dc:creator><pubDate>Fri, 17 Jul 2026 09:55:39 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/077ba3f0-664b-4b5e-ac76-c5dfbc810788.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>Cleaning Up Azure Key Vault: A Business and Developer’s Guide to Deleting, Bulk-Deleting, and Permanently Purging Secrets</strong></h2>
<p><em>Part of the series</em> <a href="https://harshtrivedii.hashnode.dev/"><em><strong>AI Cloud and Data Engineering Articles</strong></em></a><em>. Every pipeline in this series so far — SharePoint ingestion, AES key management, Snowflake-Azure connectivity — leans on Azure Key Vault to hold the secrets that make it work. This article is about what happens after those pipelines have been running for a year or two: the vault fills up, and there’s no “select all, delete” button for cleaning it out.</em></p>
<p><strong>Who this article is for:</strong> Sections 1–4 are written for anyone who owns or is accountable for a system that uses Key Vault — a business or IT lead who needs to understand the risk and approve a cleanup, not run the commands. Sections 5 onward are written for the developer or operator who will actually run them. Read the parts that match your role; nothing after Section 4 assumes you’re a business reader.</p>
<hr />
<h2><strong>Why this matters</strong></h2>
<p>Every application that isn’t hard-coding its passwords stores them somewhere safer — usually Azure Key Vault. That’s the right approach. But a vault that’s been backing an HR, Sales, or Finance system for a couple of years doesn’t stay small. It accumulates secrets the same way an inbox accumulates emails: some belong to systems that were retired last year, some were one-off credentials created for a proof-of-concept that never shipped, and some are old versions left behind by a rotation policy.</p>
<p>Modern enterprise AI systems make this challenge even bigger. Imagine an HR organization with <strong>80,000 employees</strong> building a machine learning model to predict attrition, to predict sales, to predict budget vs actuals, recommend training, or analyze workforce trends. For privacy and compliance reasons, the ML model should never receive an employee’s actual ID. Instead, the application generates a <strong>unique secret or token</strong> for every employee, stores the mapping securely in Azure Key Vault, and sends only the secret to the AI model. Once the model returns its predictions, the application looks up the secret and maps it back to the original employee ID.</p>
<p>This approach protects sensitive information and helps organizations meet privacy and governance requirements. However, it also means the Key Vault may eventually contain <strong>tens of thousands of employee-specific secrets</strong>. As employees leave the organization, a new employee joins, mergers happen, projects are retired, or an application accidentally generates duplicate or incorrect secrets because of a bug, those entries must be removed. Otherwise, the vault slowly fills with secrets that no longer serve any purpose.</p>
<p>The problem is that Azure provides <strong>no native “Select All → Delete” option</strong> for secrets. You can delete them one at a time in the Azure Portal, or you can automate the process yourself. For a vault containing only a few secrets, that’s manageable. But for enterprise workloads that may store <strong>one secret per employee, customer, device, tenant, or application integration</strong>, cleaning up thousands — or even hundreds of thousands — of secrets becomes a significant operational challenge.</p>
<p>That’s a genuine business and security concern, not just housekeeping:</p>
<ul>
<li><p><strong>Security risk</strong> — Every unnecessary secret increases the attack surface. Even if the associated employee or application no longer exists, the credential remains a protected asset that must still be governed.</p>
</li>
<li><p><strong>Operational complexity</strong> — When a vault contains tens of thousands of secrets, it becomes increasingly difficult for administrators to distinguish active secrets from obsolete ones, making troubleshooting and maintenance harder.</p>
</li>
<li><p><strong>Compliance and privacy</strong> — Regulations and internal governance policies often require organizations to remove credentials and identifiers associated with users or systems that no longer exist. Retaining unnecessary secrets can create audit findings and increase compliance risk.</p>
</li>
<li><p><strong>Application reliability</strong> — Bugs or failed deployments can accidentally create duplicate or incorrectly generated secrets. Without an efficient bulk cleanup mechanism, correcting these mistakes becomes slow, error-prone, and time-consuming.</p>
</li>
</ul>
<p>This is why understanding <strong>how to safely delete, bulk-delete, and automate secret cleanup in Azure Key Vault</strong> is an important operational capability for both enterprise teams and developers. It’s not simply about keeping a vault tidy — it’s about maintaining security, ensuring compliance, and keeping large-scale cloud applications manageable as they grow.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/11b12a6e-eba5-4eca-b118-7071b7333968.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2><strong>What Azure Key Vault actually stores</strong></h2>
<p>Azure Key Vault is Microsoft’s managed service for holding the sensitive values an application needs but should never have sitting in its source code or a config file. It manages three kinds of objects:</p>
<ul>
<li><p><strong>Secrets</strong> — arbitrary sensitive strings: database passwords, API keys, connection strings, OAuth tokens, employee secret IDs for ML models. This article is about secrets specifically.</p>
</li>
<li><p><strong>Keys</strong> — cryptographic keys used for encryption, signing, and wrapping other keys.</p>
</li>
<li><p><strong>Certificates</strong> — TLS/SSL certificates, with automated renewal support.</p>
</li>
</ul>
<p>The value of centralizing this: developers reference a secret by name and vault URL at runtime, the value itself is never in a repository that could be exposed, access is controlled through Azure identity (RBAC or access policies) rather than a shared password, and every read or write is logged for audit.</p>
<hr />
<h2><strong>Why secrets multiply so quickly</strong></h2>
<p>Secrets aren’t created just once — they’re created <strong>per user, per application, per environment, or per integration</strong>, and that adds up surprisingly fast in an enterprise.</p>
<p><strong>Consider a few common business scenarios:</strong></p>
<ul>
<li><p><strong>HR systems</strong> store employee secret tokens, payroll API keys, and benefits portal credentials. As organizations grow, new employees are onboarded, others leave, and new HR applications are integrated, causing the number of secrets to increase continuously.</p>
</li>
<li><p><strong>Sales platforms</strong> rely on CRM API tokens, quote engine credentials, and partner portal secrets. Every new customer integration, marketing campaign, or third-party service introduces additional credentials that must be securely managed.</p>
</li>
<li><p><strong>Finance applications</strong> maintain ERP connection strings, banking or SFTP credentials, and OAuth tokens for reporting platforms. Regular credential rotation, audit requirements, and temporary project environments constantly generate new secrets and new versions of existing ones.</p>
</li>
<li><p><strong>AI and Machine Learning workloads</strong> often use employee-to-token mappings, encryption keys, model access tokens, and inference API credentials. Privacy regulations frequently require organizations to replace sensitive identifiers with secure tokens before sending data to AI models, creating yet another category of secrets that must be maintained.</p>
</li>
</ul>
<p>Consider an enterprise HR analytics platform that processes data for <strong>80,000 employees</strong>. Instead of sending an employee’s actual ID to a machine learning model, the application first generates a unique token or secret, stores the mapping securely in Azure Key Vault, and sends only the token to the model. After the model produces its predictions, the application retrieves the original employee ID using that stored mapping.</p>
<p>This design protects sensitive employee information and helps satisfy privacy and compliance requirements. However, it also means the Key Vault may contain <strong>one secret for every employee</strong> — 80,000 secrets for employee mappings alone. If employees leave the organization, records must be removed, or a software bug accidentally generates incorrect or duplicate mappings, those secrets also need to be cleaned up.</p>
<p>Now combine that with thousands of API keys, connection strings, OAuth tokens, certificates, temporary proof-of-concept credentials, and secrets created by automated rotation policies. Over time, it’s easy for an Azure Key Vault to contain <strong>tens or even hundreds of thousands of secrets</strong>, many of which are no longer actively used. Managing that lifecycle efficiently becomes just as important as securing the secrets themselves.</p>
<hr />
<h2><strong>The one concept to understand before you delete anything: soft-delete vs. purge</strong></h2>
<p>This is the single most important thing to get right, for business owners and developers alike, before approving or running any cleanup.</p>
<p><strong>Deleting a secret in Key Vault does not remove it immediately.</strong> Key Vault has soft-delete behavior — think of it as a recycle bin. When you “delete” a secret, it moves into a recoverable, soft-deleted state for a retention period (7–90 days, 90 by default) and can still be restored during that window. <strong>Permanently</strong> removing it requires a second, separate action: <strong>purging</strong> it.</p>
<p><strong>Two details worth knowing before you plan a cleanup:</strong></p>
<ul>
<li><p><strong>Soft-delete is no longer optional.</strong> As of February 2025, Microsoft enabled soft-delete on all key vaults and removed the ability to turn it off — every vault has it, whether the team configured it or not. If you’re working with an older vault, you may find this behavior is already active even if no one deliberately set it up.</p>
</li>
<li><p><strong>Purge protection is a separate, optional safeguard</strong>, off by default. If it’s turned on for a vault, purging is blocked entirely until the retention period passes — not even an Owner-level account can bypass it. This is a deliberate design choice: it exists specifically so that a compromised or over-permissioned account can’t instantly and irreversibly destroy the vault’s contents.</p>
</li>
</ul>
<p>The practical takeaway: <strong>“delete” is safe and reversible for up to 90 days. “Purge” is not reversible, ever.</strong> Treat purge as a deliberate, reviewed decision — not something that happens automatically as part of routine cleanup.</p>
<hr />
<h2><strong>Which approach fits your situation?</strong></h2>
<p>Depending on the number of secrets you need to remove and how frequently cleanup is required, there are three common approaches:</p>
<ul>
<li><p><strong>Delete a single secret (Azure Portal or one command)</strong><br />Best when you only need to remove one or two known secrets. This is the simplest and safest option, with minimal risk of accidentally deleting the wrong secret.</p>
</li>
<li><p><strong>Scripted bulk delete (Azure PowerShell or Azure CLI)</strong><br />Ideal for cleaning up tens or hundreds of secrets on demand. By combining a simple loop with name-based filters (such as a prefix), you can safely remove groups of related secrets. This approach requires more care to ensure only the intended secrets are deleted.</p>
</li>
<li><p><strong>Automated Azure Function</strong><br />The best choice for enterprise environments where cleanup needs to be repeatable, scheduled, or triggered automatically. An Azure Function can filter secrets by prefix, process large batches, generate audit logs, and provide a controlled, scalable solution for ongoing secret lifecycle management. While it requires more initial setup, it offers the highest level of automation, consistency, and governance.</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/641fee70-a6b5-4418-b4c3-f0634400b1da.png" alt="" style="display:block;margin:0 auto" />

<p>The rest of this article is written for the developer or operator running these commands. If you’re a business stakeholder, the concepts covered so far should help you understand <strong>why secret cleanup matters</strong> and <strong>which approach best fits your organization’s needs</strong>. The sections that follow demonstrate how to implement each approach using Azure PowerShell, Azure CLI, and an automated Azure Function.</p>
<hr />
<h2><strong>Prerequisites</strong></h2>
<ul>
<li><p>Azure PowerShell (the <code>Az</code> module) or the Azure CLI, installed and signed in (<code>Connect-AzAccount</code> or <code>az login</code>).</p>
</li>
<li><p>An RBAC role or access policy on the vault that grants secret <strong>delete</strong> permission (and <strong>purge</strong>, if you’ll be permanently removing anything) — for example, the built-in <strong>Key Vault Secrets Officer</strong> role for delete, and <strong>Key Vault Purge Operator</strong> if purge access needs to be granted separately.</p>
</li>
<li><p>Throughout the examples below, the placeholder vault name <code>contoso-hr-kv</code> is used — replace it with your own.</p>
</li>
</ul>
<hr />
<h2><strong>cripted bulk delete (Azure PowerShell or Azure CLI)</strong></h2>
<h2><strong>Delete a single secret (soft-delete)</strong></h2>
<p>The simplest case: remove one secret you know by name. This operation performs a <strong>soft delete</strong>. The secret moves to the deleted state and remains recoverable for the vault’s configured retention period (typically up to 90 days), unless it is later purged.</p>
<p><strong>PowerShell:</strong></p>
<pre><code class="language-shell">$kvName = "contoso-hr-kv"
$secName = "employee-100234-key"
Remove-AzKeyVaultSecret -VaultName $kvName -Name $secName -Force
</code></pre>
<p><strong>Azure CLI:</strong></p>
<pre><code class="language-shell">az keyvault secret delete --vault-name contoso-hr-kv --name employee-100000-key
</code></pre>
<p>A note on the CLI command above: <strong>Note:</strong> At the time of writing, Microsoft Learn marks <code>az keyvault secret delete</code> as <strong>Deprecated</strong> in the Azure CLI reference, but the command continues to function and performs a soft delete as expected. Microsoft hasn't announced a removal date, but because the command is deprecated, it's a good idea to check the latest Azure CLI documentation before building long-lived automation around it.</p>
<p><strong>Portal (point-and-click equivalent):</strong> Open the vault → <strong>Objects → Secrets</strong> → select the secret → <strong>Delete</strong>.</p>
<hr />
<h2><strong>Bulk delete: Soft-delete every matching secret</strong></h2>
<p>Deleting a single secret is straightforward. The challenge comes when you need to remove <strong>dozens, hundreds, or even thousands of secrets</strong> — for example, after decommissioning an application, cleaning up employee-specific mappings, or removing secrets accidentally created during a failed deployment.</p>
<p><code>Get-AzKeyVaultSecret</code> retrieves the active secrets in a Key Vault. By piping the results to <code>Remove-AzKeyVaultSecret</code>, you can soft-delete every matching secret in a single operation.</p>
<pre><code class="language-shell">$kvName = "contoso-hr-kv" 
Get-AzKeyVaultSecret -VaultName $kvName | ForEach-Object { 
Write-Host "Deleting $($_.Name)" 
Remove-AzKeyVaultSecret -VaultName $kvName -Name $_.Name -Force 
}
</code></pre>
<p>If you’re deleting a large number of secrets, it’s helpful to log your progress. The following example numbers each deletion, making it easier to monitor the operation or review execution logs later.</p>
<pre><code class="language-shell">$kv = "contoso-hr-kv" 
$i = 0 

Get-AzKeyVaultSecret -VaultName $kv | ForEach-Object 
{ 
  $i++ $name = $_.Name 
  Remove-AzKeyVaultSecret -VaultName $kv -Name $name -Force 
  Write-Host ("[{0}] Deleted (soft): {1}" -f $i, $name) 
}
</code></pre>
<h2><strong>Always filter before bulk deletion</strong></h2>
<p>An unfiltered bulk delete against a shared production vault is risky. Enterprise Key Vaults often contain secrets used by multiple applications, environments, or business units. Accidentally deleting secrets that are still in use can lead to application failures or service outages.</p>
<p>Whenever possible, scope the operation to a specific naming convention or prefix. For example, if employee mapping secrets all begin with <code>employee-</code>, delete only those secrets.</p>
<pre><code class="language-shell">Get-AzKeyVaultSecret -VaultName $kv | Where-Object 
{ 
  $_.Name -like "employee-*" } | 
ForEach-Object 
{ 
  Remove-AzKeyVaultSecret -VaultName $kv -Name $_.Name -Force 
}
</code></pre>
<p>Using a consistent naming convention for secrets isn’t just good housekeeping — it enables safe, targeted automation. Prefixes such as <code>employee-</code>, <code>sales-</code>, <code>finance-</code>, or <code>ml-</code> allow administrators to confidently clean up only the intended group of secrets.</p>
<h2><strong>Azure CLI equivalent</strong></h2>
<p>If your organization standardizes on the Azure CLI, the following Bash example performs the same operation. It lists all secrets, filters those beginning with <code>employee-</code>, and soft-deletes each one.</p>
<pre><code class="language-shell">KV="contoso-hr-kv" 
for name in $(az keyvault secret list --vault-name "$KV" --query "[].name" -o tsv); do 
  if [[ "$name" == employee-* ]]; then 
    echo "Deleting $name" 
    az keyvault secret delete \ 
    --vault-name "$KV" \ 
    --name "$name" fi done
</code></pre>
<blockquote>
<p><em><strong>Best practice:</strong></em> <em>Before running any bulk delete in a production environment, first execute the listing command by itself and verify the names that match your filter. A quick review can prevent accidental deletion of active secrets.</em></p>
</blockquote>
<hr />
<h2><strong>Purging: permanent removal</strong></h2>
<p>Soft-deleted secrets still exist in the vault, in a recoverable state, and they still block reuse of that name. To actually reclaim the name and remove them for good, you purge them — and this step needs the <code>purge</code> permission granted separately from <code>delete</code>.</p>
<h2><strong>Purge a single deleted secret</strong></h2>
<p><strong>PowerShell:</strong></p>
<pre><code class="language-shell">$kvName = "contoso-hr-kv"
$secName = "sales-100000-key"
Remove-AzKeyVaultSecret -VaultName $kvName -Name $secName -InRemovedState -Force
</code></pre>
<p><strong>Azure CLI:</strong></p>
<pre><code class="language-shell">az keyvault secret purge --vault-name contoso-hr-kv --name sales-100000-key
</code></pre>
<h2><strong>Purge all deleted secrets</strong></h2>
<pre><code class="language-shell">$kvName = "contoso-hr-kv"

Get-AzKeyVaultSecret -VaultName $kvName -InRemovedState | ForEach-Object {
    Write-Host "Purging deleted secret $($_.Name)"
    Remove-AzKeyVaultSecret -VaultName $kvName -Name $_.Name -InRemovedState -Force
}
</code></pre>
<p><strong>Azure CLI</strong>, using the <code>recoveryId</code> from each deleted secret:</p>
<pre><code class="language-shell">KV="contoso-hr-kv"
az keyvault secret list-deleted --vault-name "$KV" --query "[].id" -o tsv | while read -r id; do
    echo "Purging $id"
    az keyvault secret purge --id "$id"
done
</code></pre>
<blockquote>
<p><em><strong>Purge is irreversible.</strong></em> <em>Once purged, a secret cannot be recovered under any circumstances. If purge protection is enabled on the vault, the purge command will fail — by design — until the retention window (up to 90 days) has fully elapsed. Treat purge as a deliberate, reviewed action, never something bundled automatically into a routine delete script.</em></p>
</blockquote>
<hr />
<h2><strong>Verify the results</strong></h2>
<p>Always confirm state before and after a cleanup pass.</p>
<p><strong>List the current (active) secrets:</strong></p>
<pre><code class="language-shell">Get-AzKeyVaultSecret -VaultName $kvName
</code></pre>
<p><strong>List the soft-deleted secrets that are still recoverable:</strong></p>
<pre><code class="language-shell">Get-AzKeyVaultSecret -VaultName $kvName -InRemovedState
</code></pre>
<p><strong>Azure CLI equivalents:</strong></p>
<pre><code class="language-shell">az keyvault secret list --vault-name contoso-hr-kv
az keyvault secret list-deleted --vault-name contoso-hr-kv
</code></pre>
<hr />
<h2><strong>Automated Azure Function</strong></h2>
<h2><strong>Automated cleanup with an Azure Function</strong></h2>
<p>For repeatable, scheduled, or on-demand cleanup — for example, purging only per-employee HR secrets once a quarter — an HTTP-triggered Python Azure Function is a good fit. It authenticates with a managed identity (no credentials in code), filters by a name prefix so it can never touch secrets outside its intended scope, deletes in one pass, optionally purges in a second pass, and returns a JSON summary you can review or log.</p>
<pre><code class="language-python">import json
import time
import traceback

import azure.functions as func
from azure.identity import DefaultAzureCredential
from azure.keyvault.secrets import SecretClient
from azure.core.exceptions import HttpResponseError, ResourceNotFoundError

# Replace with your Azure Key Vault URL
VAULT_URL = "https://contoso-kv.vault.azure.net/"

# Default settings
DEFAULT_PREFIX = "integration-"   # Only secrets starting with this prefix are processed
DEFAULT_LIMIT = 0                 # 0 = No limit
DEFAULT_PURGE = True              # Purge after soft-delete

# Initialize clients once per Function worker
credential = DefaultAzureCredential()
client = SecretClient(
    vault_url=VAULT_URL,
    credential=credential
)


def parse_bool(value, default):
    """Convert query/body parameter to boolean."""
    if value is None:
        return default
    return str(value).strip().lower() in ("1", "true", "yes", "y")


def main(req: func.HttpRequest) -&gt; func.HttpResponse:
    start_time = time.time()

    try:
        body = req.get_json() if req.method == "POST" else {}
    except Exception:
        body = {}

    prefix = body.get("prefix") or req.params.get("prefix") or DEFAULT_PREFIX
    purge = parse_bool(
        body.get("purge") or req.params.get("purge"),
        DEFAULT_PURGE
    )

    try:
        limit = int(
            body.get("limit")
            or req.params.get("limit")
            or DEFAULT_LIMIT
        )
    except Exception:
        limit = DEFAULT_LIMIT

    scanned = 0
    matched = 0
    deleted = 0
    purged = 0
    errors = []

    print(f"Starting cleanup")
    print(f"Prefix : {prefix}")
    print(f"Purge  : {purge}")
    print(f"Limit  : {'No Limit' if limit == 0 else limit}")

    # ------------------------------------------------------------------
    # Pass 1 - Soft Delete
    # ------------------------------------------------------------------

    try:
        for props in client.list_properties_of_secrets():

            scanned += 1
            secret_name = props.name

            if not secret_name.startswith(prefix):
                continue

            matched += 1

            try:
                print(f"Deleting {secret_name}")

                poller = client.begin_delete_secret(secret_name)
                poller.wait()

                deleted += 1

            except ResourceNotFoundError:
                pass

            except HttpResponseError as ex:
                errors.append({
                    "name": secret_name,
                    "stage": "delete",
                    "status": getattr(ex, "status_code", None)
                })

            except Exception as ex:
                errors.append({
                    "name": secret_name,
                    "stage": "delete",
                    "error": str(ex)
                })

            if limit and deleted &gt;= limit:
                break

    except Exception as ex:

        traceback.print_exc()

        return func.HttpResponse(
            json.dumps({
                "error": f"Delete failed: {str(ex)}"
            }),
            status_code=500,
            mimetype="application/json"
        )

    # ------------------------------------------------------------------
    # Pass 2 - Purge
    # ------------------------------------------------------------------

    if purge:

        try:

            for deleted_secret in client.list_deleted_secrets():

                secret_name = deleted_secret.name

                if not secret_name.startswith(prefix):
                    continue

                try:
                    print(f"Purging {secret_name}")

                    client.purge_deleted_secret(secret_name)

                    purged += 1

                except HttpResponseError as ex:

                    errors.append({
                        "name": secret_name,
                        "stage": "purge",
                        "status": getattr(ex, "status_code", None)
                    })

                except Exception as ex:

                    errors.append({
                        "name": secret_name,
                        "stage": "purge",
                        "error": str(ex)
                    })

                if limit and purged &gt;= limit:
                    break

        except Exception as ex:

            traceback.print_exc()

            return func.HttpResponse(
                json.dumps({
                    "error": f"Purge failed: {str(ex)}"
                }),
                status_code=500,
                mimetype="application/json"
            )

    elapsed = round(time.time() - start_time, 2)

    response = {
        "vault": VAULT_URL,
        "prefix": prefix,
        "purge": purge,
        "scanned": scanned,
        "matched": matched,
        "deleted": deleted,
        "purged": purged,
        "elapsed_seconds": elapsed,
        "errors": errors[:50]
    }

    return func.HttpResponse(
        json.dumps(response, indent=2),
        status_code=200,
        mimetype="application/json"
    )
</code></pre>
<p><strong>Example secret names</strong></p>
<pre><code class="language-plaintext">integration-salesforce-api
integration-workday-token
integration-servicenow-oauth
integration-sap-connection
integration-jira-webhook
integration-slack-bot
</code></pre>
<p>In the above code, change the secret name to the one that you have created with specific prefix.</p>
<h3><strong>How to call it</strong></h3>
<pre><code class="language-shell">curl "https://&lt;your-function-app&gt;.azurewebsites.net/api/cleanup?prefix=integration-&amp;limit=25&amp;purge=true"
</code></pre>
<pre><code class="language-shell"># Delete + purge only the first 25 HR secrets (a safe, incremental run)
curl "https://&lt;your-func-app&gt;.azurewebsites.net/api/cleanup?prefix=employee-&amp;limit=25&amp;purge=true"

# Soft-delete only, leave everything recoverable
curl "https://&lt;your-func-app&gt;.azurewebsites.net/api/cleanup?prefix=token-&amp;purge=false"
</code></pre>
<p>The prefix filter, the optional per-run limit, and the two-pass delete-then-purge design exist specifically so you can dry-run against a small, named batch first, review the JSON response, and only widen the scope once you trust the result.</p>
<hr />
<h2><strong>Best practices and safety checklist</strong></h2>
<ol>
<li><p><strong>List before you delete</strong> — verify names and counts first, every time.</p>
</li>
<li><p><strong>Filter by prefix or pattern</strong> — never run an unqualified “delete everything” against a shared vault.</p>
</li>
<li><p><strong>Soft-delete first, review, then purge</strong> as a separate, deliberate step — don’t chain them automatically by default.</p>
</li>
<li><p><strong>Use a per-run limit</strong> on large vaults, so you can validate results incrementally rather than committing to a single massive run.</p>
</li>
<li><p><strong>Understand purge protection</strong> — if it’s enabled on a vault, purge is blocked until the retention window ends, with no override.</p>
</li>
<li><p><strong>Prefer a managed identity</strong> (<code>DefaultAzureCredential</code>) over stored credentials for any automation that runs unattended.</p>
</li>
<li><p><strong>Keep the audit log on</strong> — Key Vault logs every delete and purge; confirm diagnostic logging is enabled if compliance requires an audit trail.</p>
</li>
</ol>
<hr />
<h2><strong>Bottom line</strong></h2>
<p>There’s no native “bulk delete secrets” button in Azure Key Vault. A short PowerShell or CLI loop covers on-demand cleanup for a handful of secrets; a prefix-scoped Azure Function turns it into a safe, repeatable, auditable operation for ongoing hygiene. Choose the level of automation that matches the size and sensitivity of the vault — and treat purge, specifically, as a decision that gets reviewed, not automated by default.</p>
<hr />
<h2><strong>References</strong></h2>
<ul>
<li><p><strong>Azure Key Vault soft-delete overview (mandatory since February 2025, retention periods) —</strong> <a href="https://learn.microsoft.com/en-us/azure/key-vault/general/soft-delete-overview">https://learn.microsoft.com/en-us/azure/key-vault/general/soft-delete-overview</a></p>
</li>
<li><p><strong>Azure Key Vault recovery overview (soft-delete vs. purge protection) —</strong> <a href="https://learn.microsoft.com/en-us/azure/key-vault/general/key-vault-recovery">https://learn.microsoft.com/en-us/azure/key-vault/general/key-vault-recovery</a></p>
</li>
<li><p><code>Remove-AzKeyVaultSecret</code> <strong>(PowerShell, including</strong> <code>-InRemovedState</code><strong>) —</strong> <a href="https://learn.microsoft.com/en-us/powershell/module/az.keyvault/remove-azkeyvaultsecret">https://learn.microsoft.com/en-us/powershell/module/az.keyvault/remove-azkeyvaultsecret</a></p>
</li>
<li><p><code>az keyvault secret</code> <strong>command reference (including the</strong> <code>delete</code> <strong>deprecation notice and</strong> <code>purge</code><strong>/</strong><code>list-deleted</code><strong>) —</strong> <a href="https://learn.microsoft.com/en-us/cli/azure/keyvault/secret">https://learn.microsoft.com/en-us/cli/azure/keyvault/secret</a></p>
</li>
<li><p><strong>Azure Key Vault Secrets client library for Python (</strong><code>begin_delete_secret</code><strong>,</strong> <code>purge_deleted_secret</code><strong>) —</strong> <a href="https://learn.microsoft.com/en-us/python/api/overview/azure/keyvault-secrets-readme">https://learn.microsoft.com/en-us/python/api/overview/azure/keyvault-secrets-readme</a></p>
<hr />
<p>#azurekeyvault #azurefunction #bulkdelete_keyvault_secrets #azurekeyvaultsecrets #purge_and_soft_delete_secrets</p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[Connecting Snowflake and Azure: Why We Do It, and a Step-by-Step Guide to Storage Integrations and Private Connectivity]]></title><description><![CDATA[Part of the series AI Cloud and Data Engineering Articles. The previous articles in this series covered encrypting sensitive documents with Azure Key Vault + AES, and building a full SharePoint → Azur]]></description><link>https://harshtrivedii.hashnode.dev/connecting-snowflake-and-azure-why-we-do-it-and-a-step-by-step-guide-to-storage-integrations-and-private-connectivity</link><guid isPermaLink="true">https://harshtrivedii.hashnode.dev/connecting-snowflake-and-azure-why-we-do-it-and-a-step-by-step-guide-to-storage-integrations-and-private-connectivity</guid><category><![CDATA[snowflake]]></category><category><![CDATA[snowflake azure connectivity]]></category><category><![CDATA[azure blob]]></category><category><![CDATA[azure-blobstorage]]></category><category><![CDATA[snowflake tutorial]]></category><dc:creator><![CDATA[Harsh Trivedi]]></dc:creator><pubDate>Mon, 13 Jul 2026 16:32:41 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/307adf84-72c3-4485-90e3-710cd87c7951.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Part of the series <a href="https://harshtrivedii.hashnode.dev/"><strong>AI Cloud and Data Engineering Articles</strong></a>. The previous articles in this series covered <a href="https://harshtrivedii.hashnode.dev/securing-enterprise-ai-pipelines-using-azure-key-vault-and-aes-encryption-for-sensitive-documents-in-snowflake-databricks-ai-workloads">encrypting sensitive documents with Azure Key Vault + AES</a>, and building a full <a href="https://harshtrivedii.hashnode.dev/building-a-secure-sharepoint-azure-blob-snowflake-document-intelligence-pipeline">SharePoint → Azure Blob → Snowflake</a> ingestion pipeline. Both of those pipelines depend on one foundational piece of plumbing that we hadn’t explicitly walked through yet: how Snowflake and Azure actually talk to each other. This article fills that gap.</p>
<hr />
<h2><strong>The Problem: Your Business Data and Your AI Platform Don’t Live in the Same Place</strong></h2>
<p>Modern enterprises rarely keep all of their data on a single platform. Customer information, financial records, invoices, contracts, operational reports, IoT data, and application logs are often generated by dozens of different business systems and stored in Azure services such as Azure Blob Storage or Azure Data Lake Storage. At the same time, organizations increasingly rely on Snowflake to power analytics, AI applications, reporting, and enterprise data sharing.</p>
<p>This creates an important business challenge. The data that drives critical business decisions is stored in one environment, while the platform responsible for analytics and AI operates in another. Moving information between these environments must be secure, reliable, and scalable. As organizations grow, manually building and maintaining separate integration pipelines for every department, application, or new data source quickly becomes expensive, difficult to govern, and operationally complex.</p>
<p>The challenge extends beyond simply loading data into Snowflake. Many enterprise workflows also require Snowflake to communicate with Azure services — for example, triggering Azure Functions to automate business processes, calling APIs exposed through Azure API Management, integrating with ERP or CRM systems, or writing results back to Azure-hosted applications. These interactions are often part of mission-critical business operations rather than simple data movement.</p>
<p>For industries such as banking, healthcare, insurance, manufacturing, and government, the question is no longer just <strong>“Can these platforms communicate?”</strong> The real question is <strong>“Can they communicate in a way that satisfies our security, compliance, governance, and regulatory requirements?”</strong> Whether data travels across the public internet or remains entirely within a private network can determine whether an architecture meets internal security standards and regulatory obligations.</p>
<hr />
<h2><strong>The Need: A Secure, Enterprise-Grade Connection Between Snowflake and Azure</strong></h2>
<p>As organizations scale their data and AI initiatives, connecting Snowflake with Azure becomes more than a technical requirement — it becomes a governance and security priority. While there are several ways to establish connectivity, enterprise organizations need a solution that is secure, easy to manage, and aligned with corporate compliance policies.</p>
<p>Traditional approaches often rely on shared credentials or access tokens to connect cloud storage with analytics platforms. Although these methods may work for small projects or proof-of-concepts, they introduce operational overhead and security risks as the number of applications, teams, and datasets grows. Credentials expire, require regular rotation, and increase the likelihood of service interruptions or human error if not managed carefully.</p>
<p>From a business perspective, security teams prefer to grant access based on verified organizational identities rather than distributing shared secrets. This allows access to be governed through Azure’s identity and access management policies, providing clear ownership, centralized administration, and a complete audit trail of who was granted access, when it was approved, and what resources can be accessed.</p>
<p>For organizations operating in highly regulated industries such as banking, healthcare, insurance, telecommunications, and government, the requirements extend even further. Sensitive business data may be required to remain within private network boundaries, ensuring that communication between Snowflake and Azure never traverses the public internet. These controls help organizations satisfy internal security standards, customer commitments, and regulatory obligations while reducing their overall security exposure.</p>
<p>To address these enterprise requirements, Snowflake provides <strong>Storage Integrations</strong>, which enable secure, identity-based access to Azure Storage without embedding long-lived credentials in application configurations. For organizations with stricter networking and compliance requirements, <strong>Azure Private Link</strong> enables private connectivity between Snowflake and Azure services, helping ensure that data traffic remains on Microsoft’s private backbone instead of the public internet.</p>
<hr />
<h2><strong>The Solution: two connectivity models, chosen by your compliance bar</strong></h2>
<p>Snowflake gives you two ways to connect to Azure Blob Storage, and the right one depends on how sensitive the data is and what your security team requires:</p>
<ol>
<li><p><strong>Storage Integration over the public network</strong> — Snowflake authenticates to your storage account as a managed Azure AD service principal (no keys, no tokens in Snowflake), but the actual data transfer still traverses the public internet (over HTTPS/TLS). This is the default, and it’s what the vast majority of Snowflake-Azure pipelines use.</p>
</li>
<li><p><strong>Storage Integration with Private Link</strong> — the same service-principal-based authentication, but the network path between Snowflake’s VNet and your storage account’s private endpoint never leaves Microsoft’s private backbone. This requires <strong>Snowflake Business Critical Edition or higher</strong>, costs more (you pay per private connectivity endpoint plus data processed), and takes noticeably more setup — but it’s the right call when a compliance requirement says “no public internet.”</p>
</li>
</ol>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/e8b95fa0-4da0-44a5-81cf-e422329a56ff.png" alt="" style="display:block;margin:0 auto" />

<p>Below is the step-by-step for both, starting with the one nearly everyone needs first.</p>
<hr />
<h2><strong>Part A: Connect Snowflake to Azure Blob Storage over the public network (Storage Integration)</strong></h2>
<p>This is the foundational setup — the same pattern used to build the external stage in <a href="https://harshtrivedii.hashnode.dev/building-a-secure-sharepoint-azure-blob-snowflake-document-intelligence-pipeline">our SharePoint → Snowflake pipeline article</a>.</p>
<h2><strong>Step 1: Create a Storage Integration in Snowflake</strong></h2>
<pre><code class="language-sql">CREATE STORAGE INTEGRATION azure_int
  TYPE = EXTERNAL_STAGE
  STORAGE_PROVIDER = 'AZURE'
  ENABLED = TRUE
  AZURE_TENANT_ID = '&lt;your_org_tenant_id&gt;'
  STORAGE_ALLOWED_LOCATIONS = ('azure://&lt;account&gt;.blob.core.windows.net/&lt;container&gt;/&lt;path&gt;/');
</code></pre>
<p>A storage integration is a Snowflake object that stores a generated Azure AD identity for your external storage, along with an explicit allow-list of locations it’s permitted to touch. It’s the mechanism that lets you avoid supplying credentials when creating stages or loading/unloading data — the credential lives in Azure identity management, not in a Snowflake stage definition.</p>
<p><code>&lt;your_tenant_id&gt;</code> is your Microsoft Entra ID (formerly Azure AD) tenant ID, found under <strong>Microsoft Entra ID → Overview → Properties → Tenant ID</strong>. A single storage integration authenticates to exactly one tenant, so every location in <code>STORAGE_ALLOWED_LOCATIONS</code> must belong to that same tenant.</p>
<h2><strong>Step 2: Retrieve the consent URL and service principal name</strong></h2>
<pre><code class="language-sql">DESC STORAGE INTEGRATION azure_int;
</code></pre>
<p><strong>Note two values from the output:</strong></p>
<p><strong>AZURE_CONSENT_URL —</strong> navigate to this URL in a browser and click Accept. This grants the Azure AD application that Snowflake created for your account a permission to request an access token on resources inside your tenant (it does not, by itself, grant access to any data — that's the next step).</p>
<p><strong>AZURE_MULTI_TENANT_APP_NAME —</strong> the name of the Snowflake service principal that Snowflake created, formatted as _. You'll search for the part before the underscore in the next step.</p>
<h2><strong>Step 3: Grant the Snowflake service principal access in the Azure Portal</strong></h2>
<ol>
<li><p>Go to <strong>Azure Portal → Storage Accounts → your storage account</strong>.</p>
</li>
<li><p>Click <strong>Access Control (IAM) → Add role assignment</strong>.</p>
</li>
<li><p>Assign one of:</p>
</li>
</ol>
<ul>
<li><p><strong>Storage Blob Data Reader</strong> — read-only access, for loading data into Snowflake.</p>
</li>
<li><p><strong>Storage Blob Data Contributor</strong> — read + write access, needed if you also want Snowflake to unload data or run <code>REMOVE</code> against files in the container.</p>
</li>
</ul>
<p>4. Under Members, search for the Snowflake service principal using the string before the underscore in AZURE_MULTI_TENANT_APP_NAME.</p>
<p>5. Click Review + assign.</p>
<p>It’s normal — not a sign something broke — for the service principal to take up to an hour (occasionally longer) to appear as searchable in the Azure Portal after you accept the consent URL, and for the role assignment itself to take a few minutes to propagate once saved.</p>
<h2><strong>Step 4: Back on Snowflake - Create the external stage</strong></h2>
<pre><code class="language-sql">LIST @my_azure_stage;
</code></pre>
<p>If the role assignment has propagated and the consent has been accepted, this returns the files sitting in that container/path of your Azure storage. If it comes back empty or errors out, the most common causes are: the role assignment hasn’t propagated yet, the consent URL wasn’t accepted, or the <code>URL</code> in the stage doesn't exactly match a location listed in <code>STORAGE_ALLOWED_LOCATIONS</code>.</p>
<p><code>SYSTEM$VALIDATE_STORAGE_INTEGRATION</code> is also worth knowing about — it runs a permissions check against the integration and a specific file path and reports back what succeeded or failed, which is faster to interpret than a generic <code>LIST</code> error while you're still debugging the setup.</p>
<hr />
<h2><strong>Part B: Private connectivity between Snowflake and an Azure Blob external stage</strong></h2>
<p>Everything in Part A still applies — you still create a storage integration and an external stage. Private connectivity adds a private network path underneath it, so traffic never crosses the public internet. <strong>This feature requires Snowflake Business Critical Edition or higher</strong>; on lower editions, <code>USE_PRIVATELINK_ENDPOINT</code> isn't available as of july 2026, and you'll need to contact Snowflake to discuss an edition upgrade before proceeding.</p>
<h2><strong>Step 1: Provision the private endpoint for Blob storage in Snowflake</strong></h2>
<pre><code class="language-sql">USE ROLE ACCOUNTADMIN;

SELECT SYSTEM$PROVISION_PRIVATELINK_ENDPOINT(
  '/subscriptions/&lt;subscription_id&gt;/resourceGroups/&lt;rg_name&gt;/providers/Microsoft.Storage/storageAccounts/&lt;storage_account&gt;',
  '&lt;storage_account&gt;.blob.core.windows.net',
  'blob'
);
</code></pre>
<p>This provisions a private endpoint inside Snowflake’s own VNet that points at your storage account. It typically takes a few minutes to complete, since Snowflake is calling Azure’s own APIs behind the scenes to create the endpoint.</p>
<p>Once it’s provisioned, go to <strong>Azure Portal → your storage account → Networking → Private endpoint connections</strong>, and <strong>approve</strong> the pending connection request from Snowflake.</p>
<h2><strong>Step 2: Check endpoint status until it’s approved</strong></h2>
<pre><code class="language-sql">SELECT SYSTEM$GET_PRIVATELINK_ENDPOINTS_INFO();
</code></pre>
<p>Poll this until the relevant entry shows <code>"status": "APPROVED"</code>. You can proceed with the next steps while waiting, but the private path won't actually carry traffic until approval goes through on the Azure side.</p>
<h2><strong>Step 3: Create the storage integration with private connectivity enabled</strong></h2>
<pre><code class="language-sql">CREATE OR REPLACE STORAGE INTEGRATION my_azure_private_int
  TYPE = EXTERNAL_STAGE
  STORAGE_PROVIDER = 'AZURE'
  AZURE_TENANT_ID = '&lt;tenant_id&gt;'
  STORAGE_ALLOWED_LOCATIONS = ('azure://&lt;account&gt;.blob.core.windows.net/&lt;container&gt;/&lt;path&gt;/')
  USE_PRIVATELINK_ENDPOINT = TRUE
  ENABLED = TRUE;
</code></pre>
<p><code>USE_PRIVATELINK_ENDPOINT = TRUE</code> is the property that tells this integration to route through the private endpoint rather than the public network. If you also need public-network access to the <em>same</em> storage account for a different use case, create a second, separate storage integration with <code>USE_PRIVATELINK_ENDPOINT = FALSE</code> — you can run both side by side.</p>
<h2><strong>Step 4: Get the consent URL and service principal name</strong></h2>
<pre><code class="language-sql">DESC STORAGE INTEGRATION my_azure_private_int;
</code></pre>
<p>Same as Part A — note <code>AZURE_CONSENT_URL</code> and <code>AZURE_MULTI_TENANT_APP_NAME</code>.</p>
<h2><strong>Step 5: Create the external stage</strong></h2>
<pre><code class="language-sql">CREATE OR REPLACE STAGE my_private_stage
  URL = 'azure://&lt;account&gt;.blob.core.windows.net/&lt;container&gt;/&lt;path&gt;/'
  STORAGE_INTEGRATION = my_azure_private_int;
</code></pre>
<p>A stage that references a storage integration with <code>USE_PRIVATELINK_ENDPOINT = TRUE</code> automatically inherits that private endpoint configuration — you don't (and can't) set the property again on the stage itself.</p>
<h2><strong>Step 6: Test the connection</strong></h2>
<pre><code class="language-sql">LIST @my_private_stage;
</code></pre>
<h2><strong>Step 7: Finish the Azure-side approvals</strong></h2>
<ol>
<li><p><strong>Approve the private endpoint</strong> — Azure Portal → Storage Account → Networking → Private endpoint connections → approve the pending Snowflake endpoint (same as Step 1, confirm it’s approved).</p>
</li>
<li><p><strong>Accept the consent URL</strong> — navigate to <code>AZURE_CONSENT_URL</code> in a browser and click Accept.</p>
</li>
<li><p><strong>Grant the role to the Snowflake service principal</strong> — Storage Account → Access Control (IAM) → Add role assignment → assign <strong>Storage Blob Data Reader</strong> (or <strong>Storage Blob Data Contributor</strong>) to the Snowflake app, searching by the string before the underscore in <code>AZURE_MULTI_TENANT_APP_NAME</code>.</p>
</li>
</ol>
<p>A couple of platform-level limits worth knowing before you plan multiple private endpoints: Snowflake currently caps outbound private connectivity at <strong>five private endpoints per account</strong>, and you can’t have more than one endpoint pointed at the same Azure subresource — plan your endpoint usage across teams accordingly, and deprovision endpoints you’re no longer using so they don’t count against the limit.</p>
<hr />
<h2><strong>Which Approach Should You Choose?</strong></h2>
<p>The right approach depends on your organization’s security requirements, compliance obligations, and business priorities.</p>
<p>For most organizations, <strong>Snowflake Storage Integration</strong> is the recommended choice. It provides a secure, identity-based connection between Snowflake and Azure Storage without relying on shared credentials or manually managed access keys. It’s simple to set up, cost-effective, and suitable for the majority of analytics, reporting, data engineering, and AI workloads.</p>
<p>If your organization operates in industries such as banking, healthcare, insurance, government, or any environment with strict security and regulatory requirements, <strong>Azure Private Link</strong> may be the better option. It keeps communication between Snowflake and Azure on Microsoft’s private network instead of the public internet, helping meet compliance requirements and reducing network exposure. The trade-off is additional configuration, the need for Snowflake Business Critical Edition, and extra networking costs.</p>
<p>The good news is that choosing one approach over the other does <strong>not</strong> change how you build your data pipelines. Whether you use a standard Storage Integration or Azure Private Link, you continue to work with the same Snowflake features — such as external stages, <code>COPY INTO</code>, Snowpipe, and stored procedures. Since the connectivity layer is independent of the data processing layer, organizations can start with a standard Storage Integration and adopt Azure Private Link later if business or compliance requirements evolve, without redesigning their existing pipelines.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/fb4e4a87-de69-4a71-82e7-77edab6cb180.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2><strong>References</strong></h2>
<ul>
<li><p>Create storage integration for Azure — <a href="https://docs.snowflake.com/en/sql-reference/sql/create-storage-integration">https://docs.snowflake.com/en/sql-reference/sql/create-storage-integration</a></p>
</li>
<li><p>Create an Azure stage — <a href="https://docs.snowflake.com/en/user-guide/data-load-azure-create-stage">https://docs.snowflake.com/en/user-guide/data-load-azure-create-stage</a></p>
</li>
<li><p>Configure an Azure container for loading data (storage integration setup, role assignment) — <a href="https://docs.snowflake.com/en/user-guide/data-load-azure-config">https://docs.snowflake.com/en/user-guide/data-load-azure-config</a></p>
</li>
<li><p>Private connectivity to external stages and Snowpipe automation for Microsoft Azure — <a href="https://docs.snowflake.com/en/user-guide/data-load-azure-private">https://docs.snowflake.com/en/user-guide/data-load-azure-private</a></p>
</li>
<li><p>Private connectivity for outbound network traffic (edition requirement, endpoint limits) — <a href="https://docs.snowflake.com/en/user-guide/private-connectivity-outbound">https://docs.snowflake.com/en/user-guide/private-connectivity-outbound</a></p>
</li>
<li><p>SYSTEM$PROVISION_PRIVATELINK_ENDPOINT — <a href="https://docs.snowflake.com/en/sql-reference/functions/system_provision_privatelink_endpoint">https://docs.snowflake.com/en/sql-reference/functions/system_provision_privatelink_endpoint</a></p>
</li>
<li><p>Snowflake editions (Business Critical Edition) — <a href="https://docs.snowflake.com/en/user-guide/intro-editions">https://docs.snowflake.com/en/user-guide/intro-editions</a></p>
<hr />
<h2><strong>What’s Next?</strong></h2>
<p>Connecting Snowflake with Azure opens the door to much more than securely moving data. It also enables organizations to automate business processes that extend beyond the Snowflake platform.</p>
<p>One common challenge many teams encounter is notifications. Snowflake provides built-in email capabilities, but they are designed to send emails only to <strong>verified users within the same Snowflake account</strong>. This works well for internal administrators and users, but it isn’t sufficient for many real-world business scenarios.</p>
<p>For example, an organization may need to notify a supplier when a contract is nearing expiration, inform a customer that a document has been processed, alert an external auditor about a compliance report, or send an email to a business partner when a data quality issue is detected. Since these recipients are typically outside the Snowflake account, Snowflake’s native email functionality cannot send notifications to them directly.</p>
<p>This is where the Snowflake–Azure integration becomes even more valuable. By securely connecting Snowflake with Azure services, you can trigger external business workflows directly from your data pipelines. A Snowflake Task or stored procedure can invoke an Azure Function App to send emails using Microsoft Graph or SendGrid, or trigger a Power Automate flow to deliver notifications through Outlook, Microsoft Teams, Slack, or other enterprise applications.</p>
<p>In the next article, we’ll build this solution step by step — showing how Snowflake can securely communicate with Azure to send notifications to anyone, whether they’re an internal employee, a customer, a vendor, or another external stakeholder.</p>
<p>#snowflake #snowflake_azure_connectivity #snowflake_private_link #snowflake_private_connectivity #snowflake_multi_tenant_app_name #storage_integration #external_stage_in_snowflake</p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[Building a Secure SharePoint → Azure Blob → Snowflake Document Intelligence Pipeline]]></title><description><![CDATA[Part of the series AI Cloud and Data Engineering Articles. In Part 1 we covered why documents should be encrypted before they ever reach cloud storage, how Azure Key Vault and Service Principals keep ]]></description><link>https://harshtrivedii.hashnode.dev/building-a-secure-sharepoint-azure-blob-snowflake-document-intelligence-pipeline</link><guid isPermaLink="true">https://harshtrivedii.hashnode.dev/building-a-secure-sharepoint-azure-blob-snowflake-document-intelligence-pipeline</guid><category><![CDATA[data pipeline]]></category><category><![CDATA[data ingestion]]></category><category><![CDATA[data-engineering]]></category><category><![CDATA[sharepoint to snowflake]]></category><category><![CDATA[secure data pipeline]]></category><dc:creator><![CDATA[Harsh Trivedi]]></dc:creator><pubDate>Sat, 11 Jul 2026 08:29:50 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/d918928b-96e9-4baf-8338-32b112972a0f.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Part of the series <a href="https://harshtrivedii.hashnode.dev/">AI Cloud and Data Engineering Articles</a>. In <a href="https://harshtrivedii.hashnode.dev/securing-enterprise-ai-pipelines-using-azure-key-vault-and-aes-encryption-for-sensitive-documents-in-snowflake-databricks-ai-workloads">Part 1</a> we covered why documents should be encrypted before they ever reach cloud storage, how Azure Key Vault and Service Principals keep the AES key separate from the encrypted data, and the in-memory decrypt pattern AI systems should follow. This article builds the full pipeline on top of that foundation — from SharePoint, through an encrypted Azure Blob landing zone, into Snowflake, where a stored procedure decrypts, parses, logs, and cleans up automatically on a schedule.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/e32df747-cfc7-49c5-843c-18c1206e0d8b.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2><strong>The Problem: valuable business knowledge is trapped in SharePoint</strong></h2>
<p>Most enterprises store their most business-critical documents in SharePoint — invoices, purchase orders, vendor contracts, legal agreements, HR records, procurement files, financial reports. These documents are rich in business context, but they sit in folders where analytics systems and AI agents can’t easily reason over them. Someone can open a contract and read it, but nothing is answering “which vendor contracts expire next quarter and which of them conflict with the payment terms on file” across ten thousand documents at once.</p>
<p>Microsoft 365 Copilot helps at the individual-productivity layer — Copilot in OneDrive can compare a handful of selected files side by side, for instance. But comparing a few documents interactively is a different problem from what enterprises actually need at scale: matching thousands of invoices against purchase orders, extracting structured fields into governed tables, building a reusable semantic search index, and joining document-derived facts with ERP or finance data. That’s an information-architecture problem, and it needs a platform, not a copilot.</p>
<hr />
<h2><strong>The Need: get documents into Snowflake without giving up control of them</strong></h2>
<p>Once the goal becomes “make SharePoint content queryable and AI-ready inside Snowflake,” a new set of requirements shows up that a simple file copy doesn’t satisfy:</p>
<ul>
<li><p><strong>The content is sensitive.</strong> Contracts, payroll data, and vendor agreements can’t sit in a landing zone in plain text, even temporarily.</p>
</li>
<li><p><strong>The encryption key can’t live next to the encrypted files.</strong> If the same breach that exposes the storage account also exposes the key, encryption bought nothing.</p>
</li>
<li><p><strong>Not every file is analytics-ready as-is.</strong> ZIP archives need to be unpacked, very long PDFs need to be split so document-AI functions can process them efficiently, and unsupported formats need to be filtered out before they clog the pipeline.</p>
</li>
<li><p><strong>The ingestion process runs unattended.</strong> Nobody wants to be the person who has to remember to run a script every morning; it needs to authenticate as an application, not as a person.</p>
</li>
<li><p><strong>Processing needs to be idempotent and auditable.</strong> The same file shouldn’t get parsed twice, and someone should be able to answer “was this document processed, and when?” without digging through logs.</p>
</li>
<li><p><strong>The landing zone shouldn’t lock you into one team’s toolchain.</strong> Other systems — Databricks, Fabric, Synapse — may want to read from the same encrypted zone later.</p>
</li>
</ul>
<p>That’s the need this article answers: an ingestion and processing pipeline that treats “get the file from SharePoint” and “make the file’s content queryable in Snowflake” as two cleanly separated, independently auditable problems, with encryption and key separation enforced at every hop.</p>
<hr />
<h2><strong>The Solution: SharePoint → encrypted Azure Blob → Snowflake stored procedure</strong></h2>
<p>The solution has three layers, each with one job:</p>
<ol>
<li><p><strong>SharePoint access</strong> — a Python ingestion pipeline authenticates to Microsoft Graph as a Service Principal, recursively walks the target document library, validates and pre-processes files (ZIP extraction, long-PDF splitting), and encrypts each one with AES-256-CBC using a key pulled from Azure Key Vault — the same pattern Part 1 walked through in depth.</p>
</li>
<li><p><strong>Transport</strong> — the encrypted files land in Azure Blob Storage, a neutral landing zone that isn’t tied to any single downstream consumer.</p>
</li>
<li><p><strong>Decrypt-and-understand</strong> — a Snowflake stored procedure decrypts each file only in memory or files are temporarily stored in internal stage, runs Snowflake Cortex’s <code>AI_PARSE_DOCUMENT</code> and <code>AI_EXTRACT</code> against the plaintext, logs/stores the result in tables, and immediately removes the decrypted copy from the stage. A Snowflake Task runs this on a schedule so the whole thing operates hands-off.</p>
</li>
</ol>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/a24eb846-9934-44de-8cb2-95057f0d8cfa.png" alt="" style="display:block;margin:0 auto" />

<p>One quick note before we get into it: Snowflake now also ships a <strong>generally available Openflow Connector for SharePoint</strong> — a managed, low-code path for ingesting SharePoint files and permissions directly into Snowflake. We haven’t covered it yet in this series, and it deserves its own detailed treatment (setup, permission model, and — importantly — how it compares to the custom pipeline below on cost, control, and flexibility). That comparison is coming in the <strong>next article</strong>. For now, this piece focuses on the custom encrypted pipeline, which remains the right choice whenever you need client-side encryption before data leaves your boundary, custom file handling, or a cloud-agnostic landing zone that other systems can also consume.</p>
<hr />
<h2><strong>Architecture at a glance</strong></h2>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/ee4cbbf4-be81-4331-89ed-7177fc011eb5.png" alt="" style="display:block;margin:0 auto" />

<pre><code class="language-plaintext">SharePoint Document Library (invoices, POs, contracts, HR docs)
        │
        ▼
Microsoft Graph API  (app-only auth via Azure Service Principal)
        │
        ▼
Python Ingestion Pipeline
  • recursive traversal + pagination
  • retry / token-refresh logic
  • file-type validation
  • optional ZIP extraction, long-PDF splitting
  • AES-256-CBC encryption (key pulled from Azure Key Vault)
        │
        ▼
Azure Blob Storage  (encrypted landing zone — AWS S3 / GCP GCS also supported)
        │
        ▼
Snowflake External Stage  (storage integration)
        │
        ▼
Snowflake Internal Stage + Stored Procedure
  • decrypt file with SnowflakeFile + AES key
  • write plaintext to a temporary internal stage
  • AI_PARSE_DOCUMENT / AI_EXTRACT
  • log to FILES_PROCESSED
  • REMOVE decrypted file from stage
        │
        ▼
Structured tables + Cortex Search index
        │
        ▼
Cortex Agent / Streamlit chatbot
</code></pre>
<p>This design keeps three concerns cleanly separated: <strong>SharePoint access</strong> (Graph API + Service Principal), <strong>transport and encryption</strong> (Python + Azure Blob), and <strong>decrypt-and-understand</strong> (Snowflake stored procedure + Cortex AI functions). Each layer can be swapped independently — a different cloud storage provider, a different ingestion runtime, a different AI extraction schema — without touching the others.</p>
<hr />
<h2><strong>Solution -</strong></h2>
<h2><strong>Step 1: Register an Azure AD app for SharePoint access</strong></h2>
<p>Because the ingestion pipeline runs unattended (no interactive user logs in), it authenticates as a <strong>Service Principal</strong> using the OAuth2 client-credentials flow — tenant ID, client ID, and a client secret or certificate. This is the standard app-only pattern for backend access to Microsoft Graph.</p>
<p><strong>In Microsoft Entra ID:</strong></p>
<ol>
<li><p><strong>App registrations → New registration.</strong> Give it a clear name, e.g. <code>sp-sharepoint-snowflake-ingestion_sps</code>.</p>
</li>
<li><p>Record the <strong>Tenant ID</strong>, <strong>Client ID</strong>, and generate a <strong>Client secret</strong>.</p>
</li>
<li><p>Under <strong>API permissions</strong>, add these Microsoft Graph <strong>application</strong> permissions:</p>
</li>
</ol>
<ul>
<li><p><code>Sites.Selected</code> — scopes the app to specific SharePoint sites only, instead of tenant-wide read access</p>
</li>
<li><p><a href="http://Files.Read"><code>Files.Read</code></a><code>.All</code> (or the narrower <code>Files.SelectedOperations.Selected</code> if you want per-file granularity)</p>
</li>
<li><p><a href="http://GroupMember.Read"><code>GroupMember.Read</code></a><code>.All</code> — only if you also need to resolve SharePoint group membership for ACLs</p>
</li>
<li><p><code>User.ReadBasic.All</code> — only if you need to resolve user display names/emails from IDs</p>
</li>
</ul>
<p>4. Grant admin consent, then grant the app access to the <strong>specific SharePoint site</strong> you’re ingesting from (not the whole tenant) — this limits blast radius if the credential ever leaks.</p>
<p><code>Sites.Selected</code> is worth calling out specifically: it's the right default for any service-to-service integration — grant the app exactly the sites it needs, nothing more.</p>
<hr />
<h2><strong>Step 2: Store configuration and secrets properly</strong></h2>
<p>Everything that isn’t code goes into environment variables or a secrets manager — never hardcoded.</p>
<pre><code class="language-plaintext"># Azure / Microsoft Graph
TENANT_ID=00000000-0000-0000-0000-000000000000
CLIENT_ID=11111111-1111-1111-1111-111111111111
CLIENT_SECRET=your-client-secret
GRAPH_URL=https://graph.microsoft.com/v1.0
SITE_ID=your-sharepoint-site-id
DRIVE_ID=your-drive-id
START_PATH=Shared Documents/Finance Documents

# Azure Blob
AZURE_BLOB_CONNECTION_STRING=DefaultEndpointsProtocol=https;AccountName=...;AccountKey=...;EndpointSuffix=core.windows.net
AZURE_CONTAINER_AND_PREFIX=enterprise-documents/encrypted/sharepoint

# Azure Key Vault
AZURE_KEY_VAULT_NAME=kv-enterprise-docs
AZURE_AES_SECRET_NAME=SHAREPOINT_AES_KEY

# Processing
PDF_MAX_PAGES_PER_PART=100
MAX_ZIP_DEPTH=5
</code></pre>
<p>In production, don’t even keep <code>CLIENT_SECRET</code> and the blob connection string in a plain <code>.env</code> on disk — pull them from Key Vault at runtime using a <strong>managed identity</strong>, the same way the pipeline pulls the AES key (see <a href="https://harshtrivedii.hashnode.dev/securing-enterprise-ai-pipelines-using-azure-key-vault-and-aes-encryption-for-sensitive-documents-in-snowflake-databricks-ai-workloads">Part 1</a> for the full Key Vault + RBAC walkthrough). The <code>.env</code> above is fine for local development only.</p>
<hr />
<h2><strong>Step 3: The Python ingestion pipeline — SharePoint → encrypted Azure Blob</strong></h2>
<p>At a high level, the script:</p>
<ol>
<li><p>Authenticates to Microsoft Graph using client-credentials flow, refreshing the token on <code>401</code> responses.</p>
</li>
<li><p>Recursively walks the configured SharePoint drive/folder, paginating through Graph API responses, with retry/backoff on transient failures.</p>
</li>
<li><p>Validates file types against an allow-list (PDF, DOCX, PPTX, images, HTML, TXT, EML, etc.), extracts ZIP archives, and optionally splits very long PDFs into smaller parts so downstream parsing stays efficient.</p>
</li>
<li><p>Fetches the AES-256 key from Azure Key Vault (never hardcoded) and encrypts each file with AES-256-CBC before it ever touches Blob Storage.</p>
</li>
<li><p>Uploads the encrypted bytes to Azure Blob, writes a small control/metadata file alongside it (source path, category, run ID), and appends to a local “already uploaded” log so re-runs don’t duplicate work.</p>
</li>
</ol>
<pre><code class="language-python">import os, io, base64, time
from datetime import datetime, timezone
import requests
from Crypto.Cipher import AES
from Crypto.Random import get_random_bytes
from azure.storage.blob import BlobServiceClient, ContentSettings
from azure.identity import DefaultAzureCredential
from azure.keyvault.secrets import SecretClient

def get_access_token(tenant_id, client_id, client_secret):
    url = f"https://login.microsoftonline.com/{tenant_id}/oauth2/v2.0/token"
    data = {
        "grant_type": "client_credentials",
        "client_id": client_id,
        "client_secret": client_secret,
        "scope": "https://graph.microsoft.com/.default",
    }
    resp = requests.post(url, data=data, timeout=60)
    resp.raise_for_status()
    return resp.json()["access_token"]

def load_aes_key(vault_name: str, secret_name: str) -&gt; bytes:
    vault_url = f"https://{vault_name}.vault.azure.net"
    credential = DefaultAzureCredential(exclude_interactive_browser_credential=True)
    client = SecretClient(vault_url=vault_url, credential=credential)
    encoded_key = client.get_secret(secret_name).value.strip()
    key = base64.b64decode(encoded_key, validate=True)
    if len(key) != 32:
        raise ValueError("AES key must be 32 bytes for AES-256.")
    return key

def aes_encrypt(file_bytes: bytes, key: bytes) -&gt; bytes:
    iv = get_random_bytes(16)
    cipher = AES.new(key, AES.MODE_CBC, iv)
    pad_len = 16 - (len(file_bytes) % 16)
    padded = file_bytes + bytes([pad_len]) * pad_len
    return iv + cipher.encrypt(padded)  # prepend IV so the decrypt step can read it back out

def upload_encrypted_blob(blob_service: BlobServiceClient, container: str,
                           blob_path: str, encrypted_bytes: bytes) -&gt; None:
    container_client = blob_service.get_container_client(container)
    container_client.upload_blob(
        name=blob_path,
        data=encrypted_bytes,
        overwrite=True,
        content_settings=ContentSettings(content_type="application/octet-stream"),
    )

def run_ingestion(config: dict) -&gt; None:
    token = get_access_token(config["tenant_id"], config["client_id"], config["client_secret"])
    aes_key = load_aes_key(config["vault_name"], config["aes_secret_name"])
    blob_service = BlobServiceClient.from_connection_string(config["blob_connection_string"])

    for file_name, file_bytes, source_path in walk_sharepoint_drive(token, config):
        encrypted = aes_encrypt(file_bytes, aes_key)
        blob_path = f"{config['prefix']}/{datetime.now(timezone.utc):%Y/%m/%d}/{file_name}.enc"
        upload_encrypted_blob(blob_service, config["container"], blob_path, encrypted)
        print(f"Uploaded {source_path} -&gt; {blob_path}")
</code></pre>
<p><code>walk_sharepoint_drive(...)</code> is the recursive Graph-API traversal function — pagination via <code>@odata.nextLink</code>, a 401-triggered token refresh, and retry-with-backoff around each request. It's mechanical and mostly boilerplate, so it's omitted here for space, but the shape is: list children of <code>DRIVE_ID</code>/<code>START_PATH</code>, recurse into folders, download file content for anything matching the supported-extension allow-list, and yield <code>(file_name, file_bytes, source_path)</code> tuples.</p>
<h2><strong>Where does this script actually run?</strong></h2>
<p>You have two solid options, and the right one depends on how “always-on” you need ingestion to be:</p>
<ul>
<li><p><strong>Local machine or company laptop, via VS Code or any IDE.</strong> Perfectly fine for development, backfills, and one-off runs. Point-and-click debugging, easy to inspect logs, no cloud deployment overhead. The downside is obvious: if the laptop is asleep, off, or on a different network, ingestion doesn’t run.</p>
</li>
<li><p><strong>Azure Function App (timer-triggered).</strong> The same script, deployed as a Python Azure Function, running serverless on a schedule (e.g., every 15–30 minutes) so new SharePoint files land in Blob without anyone’s laptop needing to be on. The Function App can use a <strong>managed identity</strong> to pull secrets from Key Vault at runtime instead of a client secret in <code>local.settings.json</code>, which is one credential fewer to rotate and protect.</p>
</li>
</ul>
<p>Most teams prototype locally in VS Code first, confirm the encryption and upload logic behaves correctly against a test SharePoint site, and then redeploy the same code as an Azure Function once it’s stable. Nothing about the script itself changes between the two — only how and where it’s triggered.</p>
<hr />
<h2><strong>Step 4: Snowflake External Stage — pointing at the Blob landing zone</strong></h2>
<p>Once encrypted files are landing in Blob, create a <strong>storage integration</strong> and an <strong>external stage</strong> in Snowflake so it can see that container:</p>
<pre><code class="language-sql">CREATE STORAGE INTEGRATION my_azure_integration
  TYPE = EXTERNAL_STAGE
  STORAGE_PROVIDER = 'AZURE'
  ENABLED = TRUE
  AZURE_TENANT_ID = '&lt;tenant-id&gt;'
  STORAGE_ALLOWED_LOCATIONS = ('azure://&lt;storage-account&gt;.blob.core.windows.net/&lt;container&gt;/&lt;prefix&gt;/');

CREATE STAGE raw_encrypted_external_stage
  URL = 'azure://&lt;storage-account&gt;.blob.core.windows.net/&lt;container&gt;/&lt;prefix&gt;/'
  STORAGE_INTEGRATION = my_azure_integration;
</code></pre>
<p><code>DESC STORAGE INTEGRATION my_azure_integration</code> returns a consent URL and an Azure AD application object ID — grant that identity <strong>Storage Blob Data Reader</strong> on the container so Snowflake can list and pull files without any long-lived key sitting in Snowflake itself.</p>
<p>From here, a scheduled <code>COPY FILES</code> (or a directory-table-driven process) moves new encrypted objects from the external stage into an <strong>internal stage(optional)</strong>, where the decrypt-and-parse stored procedure picks them up. Keeping a Snowflake-managed internal stage as the working area — rather than decrypting directly against the external stage — keeps the blast radius of the decrypted plaintext limited to Snowflake's own encrypted storage.</p>
<pre><code class="language-sql">CREATE STAGE IF NOT EXISTS raw_encrypted_stage
  ENCRYPTION = (TYPE = 'SNOWFLAKE_SSE')
  DIRECTORY = (ENABLE = TRUE);

CREATE STAGE IF NOT EXISTS decrypted_stage
  ENCRYPTION = (TYPE = 'SNOWFLAKE_SSE')
  DIRECTORY = (ENABLE = TRUE);
</code></pre>
<p>We can also directly decrypt files by placing files from external stage -&gt; decryption SP -&gt; Internal Stage.</p>
<hr />
<h2><strong>Step 5: The core piece — decrypt, parse, log, and clean up in one stored procedure</strong></h2>
<p>This is the heart of the pipeline: a Python stored procedure that</p>
<ol>
<li><p>reads the encrypted bytes off the internal/external stage(Azure blob encrypted files) using the <code>SnowflakeFile</code> class,</p>
</li>
<li><p>decrypts them with the AES key,</p>
</li>
<li><p>writes the plaintext to a temporary internal stage,</p>
</li>
<li><p>calls <code>AI_PARSE_DOCUMENT</code> (full text/layout) and <code>AI_EXTRACT</code> (structured fields) on the decrypted file,</p>
</li>
<li><p>logs the file as processed into a supporting table, and</p>
</li>
<li><p>deletes the decrypted file from the internal stage so nothing sensitive lingers.</p>
</li>
</ol>
<p>First, the supporting tables:</p>
<pre><code class="language-sql">CREATE TABLE IF NOT EXISTS files_processed (
    file_name        STRING,
    status           STRING,           -- 'parsed' | 'failed'
    error_message    STRING,
    processed_time   TIMESTAMP_NTZ DEFAULT CURRENT_TIMESTAMP()
);

CREATE TABLE IF NOT EXISTS parsed_results (
    file_name        STRING,
    parsed_content   VARIANT,   -- output of AI_PARSE_DOCUMENT
    extracted_fields VARIANT,   -- output of AI_EXTRACT
    processed_time   TIMESTAMP_NTZ DEFAULT CURRENT_TIMESTAMP()
);
</code></pre>
<p>Store the AES key as a Snowflake <strong>secret</strong> object bound to an <strong>external access integration</strong>, or fetch it through a secure UDF that calls Key Vault — either way, never inline the key value in the procedure body. The example below assumes a Snowflake secret named <code>aes_key_secret</code> holding the same base64-encoded 32-byte key that Key Vault stores.</p>
<pre><code class="language-sql">CREATE OR REPLACE PROCEDURE decrypt_and_parse_sp(encrypted_file_path STRING)
RETURNS STRING
LANGUAGE PYTHON
RUNTIME_VERSION = '3.11'
PACKAGES = ('snowflake-snowpark-python', 'pycryptodome')
HANDLER = 'run'
EXECUTE AS OWNER
AS
$$
import io
import base64
from Crypto.Cipher import AES
from snowflake.snowpark.files import SnowflakeFile


def get_aes_key(session) -&gt; bytes:
    # Pulls the base64-encoded AES-256 key from a Snowflake SECRET object.
    # The secret itself is populated out-of-band from Azure Key Vault, so the
    # key value is never typed into SQL or checked into source control.
    row = session.sql(
        "SELECT SYSTEM$GET_PRIVATE_SECRET('aes_key_secret')"
    ).collect()
    return base64.b64decode(row[0][0])


def run(session, encrypted_file_path: str) -&gt; str:
    file_name = encrypted_file_path.split('/')[-1]
    decrypted_name = file_name[:-4] if file_name.endswith('.enc') else file_name

    try:
        # 1. Read the encrypted bytes straight from the internal stage
        with SnowflakeFile.open(encrypted_file_path, 'rb') as f:
            raw = f.read()

        iv, ciphertext = raw[:16], raw[16:]
        key = get_aes_key(session)

        # 2. Decrypt (AES-256-CBC) and strip PKCS7 padding
        cipher = AES.new(key, AES.MODE_CBC, iv)
        padded = cipher.decrypt(ciphertext)
        plaintext = padded[:-padded[-1]]

        # 3. Write the plaintext to a temporary internal stage
        session.file.put_stream(
            io.BytesIO(plaintext),
            f"@decrypted_stage/{decrypted_name}",
            auto_compress=False,
            overwrite=True,
        )

        # 4. Parse with Cortex AI functions
        parsed = session.sql(f"""
            SELECT AI_PARSE_DOCUMENT(
                TO_FILE('@decrypted_stage', '{decrypted_name}'),
                {{'mode': 'LAYOUT'}}
            )
        """).collect()[0][0]

        extracted = session.sql(f"""
            SELECT AI_EXTRACT(
                file =&gt; TO_FILE('@decrypted_stage', '{decrypted_name}'),
                responseFormat =&gt; {{
                    'schema': {{
                        'type': 'object',
                        'properties': {{
                            'document_type': {{'description': 'What type of document is this (invoice, PO, contract, HR record)?', 'type': 'string'}},
                            'key_dates':     {{'description': 'Any important dates mentioned in the document', 'type': 'array'}},
                            'parties':       {{'description': 'Named organizations or individuals referenced', 'type': 'array'}}
                            'other parameters........ TCV, payment terms, bill to, ship to....etc'
                        }}
                    }}
                }}
            )
        """).collect()[0][0]

        # 5. Log the parsed content and mark the file as processed
        session.sql(
            "INSERT INTO parsed_results (file_name, parsed_content, extracted_fields) "
            "SELECT ?, PARSE_JSON(?), PARSE_JSON(?)",
            params=[decrypted_name, parsed, extracted],
        ).collect()

        session.sql(
            "INSERT INTO files_processed (file_name, status) VALUES (?, 'parsed')",
            params=[decrypted_name],
        ).collect()

        # 6. Cleanup — remove the decrypted plaintext and the original encrypted file
        session.sql(f"REMOVE @decrypted_stage/{decrypted_name}").collect()
        session.sql(f"REMOVE {encrypted_file_path}").collect()

        return f"Processed and cleaned up {decrypted_name}"

    except Exception as exc:
        session.sql(
            "INSERT INTO files_processed (file_name, status, error_message) VALUES (?, 'failed', ?)",
            params=[decrypted_name, str(exc)],
        ).collect()
        # Best effort cleanup even on failure, so plaintext never lingers
        session.sql(f"REMOVE @decrypted_stage/{decrypted_name}").collect()
        raise
$$;
</code></pre>
<p>A few things worth calling out about this procedure:</p>
<ul>
<li><p><a href="http://SnowflakeFile.open"><code>SnowflakeFile.open</code></a><code>()</code> is the documented way to stream a file — of any size — out of a stage inside a Python handler, and it works for both internal and external stages. Reading it as bytes (<code>'rb'</code>) keeps the AES decryption simple, since AES-CBC operates on raw bytes.</p>
</li>
<li><p><code>session.file.put_stream()</code> uploads the decrypted bytes directly to the internal stage from memory, without ever touching local disk inside the sandbox — the plaintext exists only as long as the procedure needs it.</p>
</li>
<li><p><code>AI_PARSE_DOCUMENT</code> returns the full text/layout of the document as JSON — this is what you want for building a Cortex Search index or doing general-purpose RAG over the document.</p>
</li>
<li><p><code>AI_EXTRACT</code> answers <em>specific</em> questions against a schema you define — this is what you want for pulling structured fields (invoice totals, contract parties, key dates) into governed tables. Use both together: parse for search, extract for structured analytics.</p>
</li>
<li><p><code>REMOVE</code> deletes the file from the stage immediately after processing, so the plaintext copy has the shortest possible lifetime. The original encrypted file is also removed once processing succeeds, so nothing double-processes on the next run.</p>
</li>
<li><p>Wrap the whole thing in a <code>try/except</code> that still logs a <code>'failed'</code> row and still attempts cleanup — you don't want a single malformed PDF to leave a decrypted file sitting in a stage indefinitely, or to silently break the pipeline for every file after it.</p>
</li>
</ul>
<hr />
<h2><strong>Step 6: Automate it with a Snowflake Task, every 1–2 hours or as per business requirements</strong></h2>
<p>The procedure above handles a single file. A thin driver procedure lists whatever hasn’t been processed yet and calls it for each new file, and a <strong>Task</strong> runs that driver on a schedule.</p>
<pre><code class="language-sql">CREATE OR REPLACE PROCEDURE process_new_files_sp()
RETURNS STRING
LANGUAGE PYTHON
RUNTIME_VERSION = '3.11'
PACKAGES = ('snowflake-snowpark-python')
HANDLER = 'run'
AS
$$
def run(session):
    new_files = session.sql("""
        SELECT relative_path
        FROM DIRECTORY(@raw_encrypted_stage)
        WHERE relative_path NOT IN (
            SELECT file_name FROM files_processed WHERE status = 'parsed'
        )
    """).collect()

    processed = 0
    for row in new_files:
        path = f"@raw_encrypted_stage/{row['RELATIVE_PATH']}"
        session.call('decrypt_and_parse_sp', path)
        processed += 1

    return f"Processed {processed} new file(s)."
$$;

-- Runs every hour
CREATE OR REPLACE TASK decrypt_and_parse_task
  WAREHOUSE = my_wh
  SCHEDULE = '60 MINUTE'
AS
  CALL process_new_files_sp();

-- Or, if you'd rather run every 2 hours:
-- SCHEDULE = '120 MINUTE'
-- Or with cron syntax, e.g. on the hour, every 2 hours, UTC:
-- SCHEDULE = 'USING CRON 0 */2 * * * UTC'

-- Tasks are created in a SUSPENDED state — you must explicitly resume them
ALTER TASK decrypt_and_parse_task RESUME;
</code></pre>
<p><strong>A couple of operational notes:</strong></p>
<p>DIRECTORY(@raw_encrypted_stage) requires the stage to have been created with DIRECTORY = (ENABLE = TRUE), which gives you a queryable directory table of everything sitting in the stage — handy for exactly this "what's new" pattern. The SCHEDULE parameter accepts either a fixed interval ('60 MINUTE', '120 MINUTE') or a USING CRON expression if you need it to run at a specific time of day rather than a rolling interval. Every newly created task starts suspended — ALTER TASK ... RESUME is what actually turns it on. It's an easy step to forget and then wonder why nothing is happening. SUSPEND_TASK_AFTER_NUM_FAILURES is worth setting so a persistently broken file (or an expired credential) doesn't spin the task forever without anyone noticing — check INFORMATION_SCHEMA.TASK_HISTORY() periodically, or alert off it.</p>
<pre><code class="language-sql">CREATE STORAGE INTEGRATION my_s3_integration
  TYPE = EXTERNAL_STAGE
  STORAGE_PROVIDER = 'S3'
  ENABLED = TRUE
  STORAGE_AWS_ROLE_ARN = 'arn:aws:iam::123456789012:role/snowflake-role'
  STORAGE_ALLOWED_LOCATIONS = ('s3://my-bucket/encrypted-docs/');

CREATE STAGE raw_encrypted_external_stage
  URL = 's3://my-bucket/encrypted-docs/'
  STORAGE_INTEGRATION = my_s3_integration;
</code></pre>
<p>Google Cloud Storage:</p>
<pre><code class="language-sql">CREATE STORAGE INTEGRATION my_gcs_integration
  TYPE = EXTERNAL_STAGE
  STORAGE_PROVIDER = 'GCS'
  ENABLED = TRUE
  STORAGE_ALLOWED_LOCATIONS = ('gcs://my-bucket/encrypted-docs/');

CREATE STAGE raw_encrypted_external_stage
  URL = 'gcs://my-bucket/encrypted-docs/'
  STORAGE_INTEGRATION = my_gcs_integration;
</code></pre>
<p>If you’re on AWS or GCP already, the AES key can live in AWS Secrets Manager or Google Secret Manager instead of Azure Key Vault — same principle (key separated from encrypted data, fetched only at decrypt time), different vendor.</p>
<hr />
<h2><strong>Recap: what this pipeline gives you</strong></h2>
<ul>
<li><p><strong>Encryption in transit and at rest</strong> from the moment a file leaves SharePoint until the instant before it’s parsed.</p>
</li>
<li><p><strong>Key separation</strong> — the AES key never lives next to the encrypted files, in Blob, in an external stage, or in source control.</p>
</li>
<li><p><strong>Least-lifetime plaintext</strong> — the decrypted file exists in an internal Snowflake stage only for the duration of parsing, and is removed immediately after (and even on failure).</p>
</li>
<li><p><strong>Auditability</strong> — <code>files_processed</code> gives you a record of every file, when it ran, and whether it succeeded, without re-processing anything twice.</p>
</li>
<li><p><strong>Automation</strong> — a Snowflake Task keeps the whole thing running hands-off every hour (or two), and the Python ingestion side can run from a developer’s laptop during testing or as an unattended Azure Function in production.</p>
</li>
</ul>
<p><strong>Portability</strong> — swap Azure Blob for S3 or GCS, or swap Key Vault for another secrets manager, without touching the Snowflake-side logic at all.</p>
<hr />
<h2>What’s next</h2>
<ul>
<li><p><code>AI_PARSE_DOCUMENT</code> and <code>AI_EXTRACT</code> hand back JSON — a VARIANT sitting in parsed_results. That's a good stopping point for this article, but it's not yet something a business user or a downstream table can consume directly, and there's a real amount of design work between "we have JSON" and "we have a governed, queryable dataset."</p>
<p>In the next article, we’ll take a close look at the Openflow Connector for SharePoint — how to set it up, its permission model (Sites.Selected, document ACLs), and a head-to-head comparison against the custom pipeline built here: where the managed connector wins on speed of setup, and where the custom pipeline still wins on encryption control, custom file handling, and cross-platform reuse.</p>
<p><strong>After that, we’ll come back to this pipeline and go deep on turning that raw JSON into something an AI agent and a business user can both actually use:</strong></p>
<ul>
<li><p><strong>Structured side (</strong><code>AI_EXTRACT</code> <strong>output):</strong> landing the raw JSON in a staging table as-is, then a clean/cast step that pulls fields out of the <code>VARIANT</code> into proper typed columns (dates as <code>DATE</code>, amounts as <code>NUMBER</code>, not strings) in a target table — and why you want that staging layer instead of casting inline during extraction. From there, building views that join the extracted document fields against your existing business unit, customer, and contract-master tables, so "which vendor contracts expire next quarter" is a normal SQL query, not a document search.</p>
</li>
<li><p><strong>Unstructured side (</strong><code>AI_PARSE_DOCUMENT</code> <strong>output):</strong> why the parsed JSON needs to be <strong>flattened</strong> before it's usable, how to <strong>chunk</strong> the flattened text sensibly (page-aware, section-aware, or token-based chunking, and the tradeoffs between them), and how to build a <strong>Cortex Search</strong> service on top of the chunked content.</p>
</li>
<li><p><strong>Putting it together:</strong> a full walkthrough of standing up a <strong>Cortex Agent</strong> that reasons across both the structured tables and the Cortex Search index in one conversation, plus the <strong>end-user interface</strong> on top of it — whether that’s a Streamlit-in-Snowflake app or an external Angular front end calling the agent through an API.</p>
</li>
</ul>
</li>
</ul>
<p>That’s the article where this pipeline stops being “documents in a stage” and starts being an actual AI application people can talk to.</p>
<p>#snowflake #dataengineering #sharepointtosnowflake #graphapi #AIAGENT #agentic_pipeline #sharepoint_to_snowflake #dataingestion</p>
]]></content:encoded></item><item><title><![CDATA[Moving from Snowflake Streamlit Warehouse Runtime to Snowpark Container Services: Restricted Caller Rights, Installing Third-Party Packages and Building Enterprise AI Applications]]></title><description><![CDATA[Enterprises are increasingly moving beyond static BI dashboards (Power BI, Tableau, etc.) toward AI-driven data exploration. Employees from HR to Finance now expect conversational access to data: for ]]></description><link>https://harshtrivedii.hashnode.dev/moving-from-snowflake-streamlit-warehouse-runtime-to-snowpark-container-services-restricted-caller-rights-installing-third-party-packages-and-building-enterprise-ai-applications</link><guid isPermaLink="true">https://harshtrivedii.hashnode.dev/moving-from-snowflake-streamlit-warehouse-runtime-to-snowpark-container-services-restricted-caller-rights-installing-third-party-packages-and-building-enterprise-ai-applications</guid><category><![CDATA[snowflake]]></category><category><![CDATA[snowflake snowpark]]></category><category><![CDATA[snowflake streamlit]]></category><category><![CDATA[cortex agent]]></category><category><![CDATA[snowpark container service]]></category><category><![CDATA[ai-agent]]></category><category><![CDATA[Enterprise AI]]></category><dc:creator><![CDATA[Harsh Trivedi]]></dc:creator><pubDate>Sat, 04 Jul 2026 13:48:57 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/e9c526f4-1721-4b04-bcd3-94c4d993b3ce.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Enterprises are increasingly moving beyond static BI dashboards (Power BI, Tableau, etc.) toward AI-driven data exploration. Employees from HR to Finance now expect conversational access to data: for example, an HR analyst might ask an AI agent about <em>attrition rates, background verification status, or employee headcount trends</em>, while finance queries <em>revenue, PAT, or customer billing summaries</em> conversationally.</p>
<p>Streamlit in Snowflake’s <em>Warehouse Runtime</em> was a natural first step for building such interfaces (with dashboards on one page and an AI chat on another). However, Warehouse Runtime has key limitations: it runs all code with the <em>app owner’s</em> privileges (no restricted caller’s rights) and limits Python packages to Snowflake’s curated Conda channel. This makes it nearly impossible to enforce per-user row-level security (RLS) or to install the rich set of third-party libraries (Streamlit releases, AI/ML packages, Azure SDKs, etc.) modern applications need.</p>
<p>To address this, we migrate our Streamlit app to <strong>Snowpark Container Services (SPCS)</strong> — Snowflake’s container-based runtime for Streamlit (and native apps) — which supports custom Python environments and <em>restricted caller’s rights</em>.</p>
<p>This guide uses the latest Snowflake documentation (as of mid-2026) and flags any preview/region dependencies.</p>
<hr />
<h2><strong>From Dashboards to AI Agents: The Enterprise Motivation</strong></h2>
<p>For years, BI tools (Power BI, Tableau, Qlik, Snowsight, etc.) dominated enterprise analytics. HR teams tracked attrition rates, background-check (BGV) compliance, hiring funnel metrics, etc., on dashboards. Finance teams reviewed revenue, PAT, cost breakdowns, and forecasts. These reports are often based on the very data <em>already in Snowflake</em>. If employees can query dashboards, why not <em>ask</em> an AI agent?</p>
<p>A modern enterprise AI assistant might combine traditional analytics and natural-language questioning. For example, imagine a Streamlit app with multiple tabs: one showing key performance indicator (KPI) visuals, and another hosting a Cortex AI agent for chat. An HR manager could open the app, authenticate via SSO, and ask “<strong>Show me attrition trends over the last year</strong>” or “<strong>How many candidates are pending background verification (BGV)</strong>.” A finance user might ask about <strong>“Quarterly revenue by region”</strong> or <strong>“Top 10 customers by PAT”</strong> and get answers in chat form. Behind the scenes, these agents use Snowflake data and AI models (cortex agents), but the user experience is seamless. This is a powerful paradigm shift: from static dashboards to interactive AI-powered data assistants.</p>
<p>To build such apps, developers often start with <strong>Streamlit in Snowflake</strong> (easy deployment, direct Snowflake access, Snowsight integration). However, the Warehouse Runtime quickly becomes a hindrance for enterprise needs.</p>
<hr />
<h2><strong>Limitations of Streamlit on Warehouse Runtime</strong></h2>
<p><strong>Warehouse Runtime</strong> runs a <em>new, isolated instance of the Streamlit app for each viewer</em>.</p>
<p>This ensures fresh compute per user, but it has significant constraints:</p>
<ul>
<li><p><strong>Owner’s Rights Only:</strong> <em>Warehouse runtime always executes queries with the app owner’s privileges.</em> There is no concept of restricted caller’s rights here. In practice, this means if Rakesh (an HR user with HR Role) and Harsh (a Finance user with Finance Role) use the app, all Snowflake queries run under the single “<strong>owner</strong>” role that deployed the app. HR User Rakesh with HR role will still able to see Finance Data, because the queries run under the <strong>OWNER’s</strong> role and not on <strong>CALLER’s</strong> role(HR role). The app cannot natively <em>act on behalf of Rakesh or Harsh</em>. In particular, if you have any <strong>Row Access Policies (RLS)</strong> that use <code>CURRENT_ROLE()</code>, they will see the owner’s role, not the viewer’s role. This breaks security: an HR user might see finance data and vice versa. (Snowflake docs explicitly warn that “Warehouse-runtime apps using <code>CURRENT_ROLE()</code> in row access policies will always return the app owner’s role, not the viewer’s role”.)</p>
</li>
<li><p><strong>Limited Packages:</strong> In Warehouse runtime, Python packages are limited to what Snowflake provides via a built-in Anaconda/Conda channel. You can only include dependencies by specifying an <code>environment.yml</code> that references the Snowflake Anaconda repository. Many popular packages (like recent Streamlit versions, Azure SDKs, LangChain, etc.) may not be available or are outdated. For example, Warehouse apps only support Streamlit versions up to ~1.22, whereas enterprise developers often want the latest 1.5x versions for new features. The <code>environment.yml</code> method is also clunky to update.</p>
</li>
<li><p><strong>No Persistent Server or Caching:</strong> Each viewer has a separate app instance. This increases resource usage and latency (each page view starts a new Python process). It also means caching (e.g. <code>@st.cache</code>) does not share data between users. The app can feel less responsive for multiple users.</p>
</li>
<li><p><strong>Network &amp; API Access:</strong> By default, Warehouse runtime has no internet access (except Snowflake). You can’t call external APIs or pip-install at runtime. Even calling an external API requires special support (not natively possible for a warehouse-streamlit app). This severely limits using AI APIs (OpenAI, Azure, etc.) or pulling data from external services.</p>
</li>
</ul>
<p>In summary, <strong>Warehouse Runtime is great for rapid prototyping and simple dashboards</strong>, but it cannot meet enterprise requirements for security, extensibility, and performance.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/123e8e27-061e-4815-93c3-1a0279706c8e.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2>Why Row Access Policies Break Under Owner’s Rights</h2>
<p>A concrete issue is row-level security (RLS) based on user roles. Suppose we have a table SALES and a row access policy that uses CURRENT_ROLE() to filter rows to only those in the user’s business unit. For example:</p>
<pre><code class="language-sql">CREATE OR REPLACE ROW ACCESS POLICY rls_sales AS ( (CURRENT_ROLE() = 'FINANCE' AND department = 'Finance') OR (CURRENT_ROLE() = 'HR' AND department = 'HR') ); 

ALTER TABLE SALES ADD ROW ACCESS POLICY rls_sales ON (department);
</code></pre>
<p>If the Streamlit app runs under owner’s rights with a fixed role (say the “app_service_role”), then <code>CURRENT_ROLE()</code> will always be <code>app_service_role</code>, no matter who is logged in. Thus an HR user or Finance user sees <em>the same filtered data</em> – likely none or all, depending on the policy logic. The user’s own Snowflake role is ignored. This is confirmed in Snowflake’s security docs: <em>“Warehouse-runtime apps using</em> <code>CURRENT_ROLE()</code> <em>in row access policies will always return the app owner’s role, not the viewer’s role”</em>.</p>
<p>Therefore, <strong>any fine-grained access control (by role, by user) fails under owner’s rights</strong>. The only workaround is to implement your own logic in the app (e.g. passing the Snowflake user token in queries), but that is fragile and shifts security checks to application code.</p>
<hr />
<h2><strong>Why Container Runtime (SPCS) Solves These Issues</strong></h2>
<p>Snowflake’s <strong>Container Runtime</strong> (also known as Snowpark Container Services or “Compute Pool for Streamlit”) was designed to address these limitations. In Container mode:</p>
<ul>
<li><p><strong>Restricted Caller’s Rights:</strong> Container runtimes support <em>restricted caller’s rights</em>, a model where an app can run with the <em>caller’s (viewer’s) Snowflake privileges</em>, subject to explicit grants. In other words, if Rakesh opens the app, the Snowflake queries can execute as Rakesh(with his assigned role) rather than as the owner’s role. Crucially, the developer must explicitly ask for each privilege (e.g. SELECT on a table) via <code>GRANT CALLER ...</code> commands. This prevents privilege escalation while enabling true user-based filtering. (Note: Restricted caller’s rights was a preview feature, but as of Aug 2025 it is GA for native apps/SPCS and fully supported in Streamlit container mode.)</p>
</li>
<li><p><strong>Shared Persistent Server:</strong> In container mode, one instance of the app runs and all viewers share it. This means faster connections for new viewers (the app is already running) and support for caching results across users (when appropriate).</p>
</li>
<li><p><strong>Unlimited Packages:</strong> The app runs in a Snowpark-managed Docker container where you control the image. You can use <em>PyPI</em> (and any internet-accessible repository) to install any Python package. There’s no Conda limitation. This unlocks the full Python ecosystem (e.g. the latest Streamlit, <code>langchain</code>, <code>openai</code>, <code>azure-identity</code>, <code>sentence-transformers</code>, etc.). It also allows including OS-level libraries and binaries if needed.</p>
</li>
<li><p><strong>Network Access via External Integrations:</strong> Container apps can access the internet, but you must explicitly allow it via an <em>External Access Integration (EAI)</em>. This means you define which network endpoints (PyPI, REST APIs, etc.) the app can reach. This controlled egress is much more flexible than Warehouse runtime, while still meeting security requirements.</p>
</li>
<li><p><strong>Flexible Config:</strong> Container apps can run any Streamlit version (including streamlit-nightly) and any Python 3.11 packages. You can configure ports, compute pool size, autosuspend, etc., via DDL (<code>CREATE/ALTER SERVICE</code>) or the Snowflake UI.</p>
</li>
</ul>
<p>Because of these advantages, <strong>enterprise AI apps on Snowflake should use Container Runtime</strong> once they move beyond proof-of-concept. The rest of this article explains the key steps to make this transition: setting up restricted caller’s rights, configuring package access, and building the container environment.</p>
<hr />
<h2><strong>Enabling Restricted Caller’s Rights</strong></h2>
<p>To run queries under the <strong>caller’s credentials</strong>, we use Snowflake’s <em>restricted caller’s rights</em> model. Conceptually, this changes the execution flow:</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/d7bf6939-c8e7-4064-b0e9-c0f1f6cbc796.png" alt="" style="display:block;margin:0 auto" />

<p>Under the hood, Snowflake restricts the app so it can only use privileges explicitly granted via GRANT CALLER ... (also called caller grants).</p>
<p>In this architecture, your app code runs in a container, but it queries Snowflake using the viewer’s identity.</p>
<p>The App Owner Role (APP_OWNER_ROLE): Owns the container/app and has full CRUD privileges on the application objects (the container itself, the stage, tables, views etc.). The Business Roles (HR_ROLE, FINANCE_ROLE): These roles have read-only privileges on the underlying datasets (Tables, Views, Semantic Views, Stored Procedures, UDFs, Stages, Cortex Analyst/Search/Agents). The Enforcement: By granting “Caller Rights,” you tell Snowflake: “When a user (e.g., Rakesh with HR_ROLE) connects via the app, the app can only execute queries using the privileges that Rakesh already possesses."</p>
<p><strong>Here’s how to set it up:</strong></p>
<ol>
<li><strong>Use Container Runtime:</strong> Ensure your Streamlit app is created or migrated to container runtime (Snowflake calls it “FROM” mode in CREATE STREAMLIT). Restricted caller’s rights <em>require</em> container runtime. In warehouse runtime, any attempt to use <code>st.connection("snowflake-callers-rights")</code> will fail.</li>
</ol>
<p><strong>In Snowflake container runtime streamlit file add-</strong></p>
<p><code>st.connection("snowflake-callers-rights")</code></p>
<pre><code class="language-python">import streamlit as st

# For enabling Snowflake caller rights
conn = st.connection("snowflake-callers-rights")

# This query will now respect RLS policies tied to CURRENT_ROLE()
df = conn.query("SELECT * FROM DATA_DB.PUBLIC.HR_DATA")
st.write(df)
</code></pre>
<ol>
<li><p><strong>Streamlit Version:</strong> Use Streamlit library version &gt;= 1.53.0. Older versions do not support the caller’s rights feature.</p>
</li>
<li><p><strong>Grant CALLER privileges:</strong> A Snowflake admin (or higher) must grant caller privileges on each object the app needs.</p>
</li>
</ol>
<p><strong>NOTE: The caller privileges will be granted only to the app owner role.</strong></p>
<p>To enable this, an Administrator (or a role with MANAGE CALLER GRANTS) must perform the grants. You can either perform these manually or grant the app owner the power to manage their own caller grants.</p>
<p>The admin can grant the manager caller grants to the owner role, this lets the app owner grant caller rights on underlying objects themselves:</p>
<pre><code class="language-sql">-- 1. Enable self-management of caller grants for the App Owner
GRANT MANAGE CALLER GRANTS ON ACCOUNT TO ROLE APP_OWNER_ROLE;

-- 2. Grant Caller Usage/Select on required Databases and Tables
-- HR Data
GRANT CALLER USAGE ON DATABASE hr_db TO ROLE APP_OWNER_ROLE;
GRANT INHERITED CALLER USAGE ON ALL SCHEMAS IN DATABASE hr_db TO ROLE APP_OWNER_ROLE;
GRANT CALLER SELECT ON TABLE hr_db.hr_schema.employees TO ROLE APP_OWNER_ROLE;

-- Finance Data
GRANT CALLER USAGE ON DATABASE finance_db TO ROLE APP_OWNER_ROLE;
GRANT INHERITED CALLER USAGE ON ALL SCHEMAS IN DATABASE finance_db TO ROLE APP_OWNER_ROLE;
GRANT CALLER SELECT ON TABLE finance_db.finance_schema.revenue TO ROLE APP_OWNER_ROLE;

-- Sales Data
GRANT CALLER USAGE ON DATABASE sales_db TO ROLE APP_OWNER_ROLE;
GRANT INHERITED CALLER USAGE ON ALL SCHEMAS IN DATABASE sales_db TO ROLE APP_OWNER_ROLE;
GRANT CALLER SELECT ON TABLE sales_db.sales_schema.pipeline TO ROLE APP_OWNER_ROLE;

-- BGV Data
GRANT CALLER USAGE ON DATABASE bgv_db TO ROLE APP_OWNER_ROLE;
GRANT INHERITED CALLER USAGE ON ALL SCHEMAS IN DATABASE bgv_db TO ROLE APP_OWNER_ROLE;
GRANT CALLER SELECT ON TABLE bgv_db.bgv_schema.bgv_data TO ROLE APP_OWNER_ROLE;

-- Operations Data
GRANT CALLER USAGE ON DATABASE ops_db TO ROLE APP_OWNER_ROLE;
GRANT INHERITED CALLER USAGE ON ALL SCHEMAS IN DATABASE ops_db TO ROLE APP_OWNER_ROLE;
GRANT CALLER SELECT ON TABLE ops_db.ops_schema.business_ops TO ROLE APP_OWNER_ROLE;

-- Semantic Views &amp; Stages
GRANT CALLER SELECT ON VIEW DATA_DB.PUBLIC.SEMANTIC_VIEW TO ROLE APP_OWNER_ROLE;
GRANT CALLER READ ON STAGE DATA_DB.PUBLIC.DOC_STAGE TO ROLE APP_OWNER_ROLE;

-- Cortex Search, Agents, and Snowflake Functions
GRANT CALLER USAGE ON CORTEX SEARCH SERVICE DATA_DB.PUBLIC.MY_SEARCH_SERVICE TO ROLE APP_OWNER_ROLE;
GRANT CALLER USAGE ON DATABASE SNOWFLAKE TO ROLE APP_OWNER_ROLE;
GRANT INHERITED CALLER USAGE ON ALL SCHEMAS IN DATABASE SNOWFLAKE TO ROLE APP_OWNER_ROLE;
GRANT INHERITED CALLER USAGE ON ALL FUNCTIONS IN DATABASE SNOWFLAKE TO ROLE APP_OWNER_ROLE;

-- Intelligence/Agents DB
GRANT CALLER USAGE ON DATABASE SNOWFLAKE_INTELLIGENCE TO ROLE APP_OWNER_ROLE;
GRANT CALLER USAGE ON SCHEMA SNOWFLAKE_INTELLIGENCE.AGENTS TO ROLE APP_OWNER_ROLE;
</code></pre>
<p>Each caller grant allows the app to use that privilege only if the viewer has it. The database and schema grants use INHERITED CALLER USAGE (to apply to all current/future schemas or objects). We also grant CALLER rights on the SNOWFLAKE system databases so Snowflake’s functions and UDFs will run under the caller’s context. Similarly, if the app calls a semantic view or stored procedure, grant CALLER SELECT or CALLER EXECUTE on those objects. For example, for a semantic view corp_db.corp_schema.kpi_view, use GRANT CALLER SELECT ON VIEW corp_db.corp_schema.kpi_view TO ROLE APP_OWNER_ROLE. For an internal stage or file format, use GRANT INHERITED CALLER USAGE,READ ON ALL STAGES IN SCHEMA ops_db.ops_schema</p>
<p><strong>Make sure the admin has provided the cortex_user role and cortex_agent_user role to the role.</strong></p>
<pre><code class="language-sql">GRANT DATABASE ROLE SNOWFLAKE.CORTEX_USER TO APP_OWNER_ROLE;
GRANT DATABASE ROLE SNOWFLAKE.CORTEX_AGENT_USER TO APP_OWNER_ROLE;
</code></pre>
<p><strong>4.Specifying Execution Rights in</strong> <code>streamlit_spec.yaml</code></p>
<p>To fully enable restricted caller’s rights for the container application, you must define the <code>executeAsCaller</code> capability in your <code>streamlit_spec.yaml</code> file within your project folder.</p>
<pre><code class="language-yaml">spec:
  containers:
    - name: streamlit-app
      image: &lt;your_image_path&gt;
      env:
        STREAMLIT_SERVER_PORT: 8080
capabilities:
  securityContext:
    executeAsCaller: true
</code></pre>
<p><strong>5.Critical Requirement: Default Role Alignment</strong></p>
<p>Even with the correct CALLER grants and executeAsCaller: true configuration, the restricted caller’s rights model relies on the user's default role.</p>
<ul>
<li><p><strong>Role Context:</strong> Snowflake’s restricted-caller connections resolve permissions based on the user’s DEFAULT_ROLE rather than the role currently selected in the Snowsight UI at the time of access.</p>
</li>
<li><p><strong>Enforcement:</strong> If a user (e.g., Rakesh) logs into the Streamlit app, the query will be executed using his DEFAULT_ROLE. If his DEFAULT_ROLE is set to PUBLIC (or any role other than HR_ROLE), the RLS policy—which checks CURRENT_ROLE()—will fail to authorize the query, resulting in "no data" or an "insufficient privileges" error.</p>
</li>
<li><p><strong>Best Practice:</strong> Ensure all the end users of this app have their DEFAULT_ROLE correctly configured to match the business role required to view the data for their specific department.</p>
</li>
</ul>
<p>By enforcing these five steps, you ensure that the application is not only functional but also compliant with strict enterprise data governance, preventing “privilege escalation” where the app owner’s rights would otherwise accidentally expose sensitive data to unauthorized users.</p>
<p>Throughout this setup, the app is still “owned” by app_owner_role, but its queries run as the viewer’s role. This is the essence of restricted caller’s rights. (The name highlights: the app is restricted to use only those caller privileges explicitly granted, rather than all of the caller’s powers.)</p>
<p><strong>Now you are all set to use caller rights. This setup will help you to restrict user’s role to access only the data they are allowed to see.</strong></p>
<p><strong>QUICK TEST:</strong></p>
<p><code>Change your default role to HR_ROLE and test the RLS on the streamlit container app — it will only fetch the data HR_ROLE is allowed to see.</code></p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/2e6563d5-f92a-414b-a339-a4662e141966.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2><strong>Major Consideration When Using Cortex Search with Restricted Caller Rights</strong></h2>
<p>One important point that isn’t immediately obvious is how <strong>Cortex Search</strong> behaves in an enterprise application.</p>
<p>A Cortex Search Service builds and maintains vector embeddings over the source table to provide semantic search capabilities. However, simply enabling <strong>Restricted Caller Rights</strong> on your Streamlit application does <strong>not</strong> automatically guarantee that every Cortex Search request is evaluated using the logged-in user’s security context.</p>
<p>In most enterprise applications, document access is controlled using <strong>Row Access Policies</strong>, business-unit mappings, customer-level authorization, or department-specific permissions. These policies often rely on functions such as <code>CURRENT_ROLE()</code> or <code>CURRENT_USER()</code> to determine which records a user is allowed to access.</p>
<p>If your Agent invokes the Cortex Search Service directly, you may not get the desired caller-context evaluation for those security policies.</p>
<h2><strong>Solution:</strong></h2>
<blockquote>
<p><em>Your agent definition should call the SP, that has the cortex search service, instead of calling the search service directly.</em></p>
</blockquote>
<p>To bridge this gap, implement a <strong>secure intermediary layer</strong> — such as a Stored Procedure or a dedicated Python function — that acts as the gatekeeper between your AI Agent and the Cortex Search Service. Because this intermediary is invoked by your Streamlit app running with <strong>Restricted Caller’s Rights</strong>, it executes using the logged-in user’s security context rather than the app owner’s.</p>
<p><strong>Your workflow should be:</strong></p>
<ol>
<li><p><strong>User Request</strong>: The user initiates a query via the Streamlit interface.</p>
</li>
<li><p><strong>Secure Intermediary</strong>: The Agent decides to use cortex analyst or search, based on instructions it calls the intermediary (Stored Procedure/Function) for detailed data or summary that is in tables on which search service is made.</p>
</li>
<li><p><strong>Context-Aware Execution</strong>: The intermediary executes the search under the <code>CURRENT_ROLE()</code> of the authenticated user.</p>
</li>
<li><p><strong>RLS Enforcement</strong>: Snowflake’s native Row Access Policies automatically evaluate the user’s role against the indexed documents.</p>
</li>
<li><p><strong>Filtered Result</strong>: Only the subset of documents authorized for that specific user is retrieved and passed back to the Agent for processing.</p>
</li>
</ol>
<hr />
<h2><strong>Network Rules and External Access Integration for Packages</strong></h2>
<p>To install third-party Python packages in SPCS Streamlit apps, you must grant the app internet access <strong>at build/deploy time</strong>. Snowflake enforces this via <em>network rules</em> and <em>external access integrations (EAI)</em>.</p>
<h2><strong>Creating Network Rules</strong></h2>
<p>A <strong>network rule</strong> specifies an external host (and port) that the container is allowed to reach. For PyPI (Python package index), we need to allow access to <a href="http://pypi.org"><code>pypi.org</code></a> and its CDN domain. You can either use Snowflake’s <em>managed PyPI rule</em> or create your own. For illustration, here’s how to create one:</p>
<pre><code class="language-sql">-- Example: allow PyPI domains (HTTPS default port 443)
CREATE OR REPLACE NETWORK RULE pypi_network
  MODE = EGRESS
  TYPE = HOST_PORT
  VALUE_LIST = ('pypi.org:443', 'files.pythonhosted.org:443');
</code></pre>
<p>This VALUE_LIST covers the main PyPI endpoints. (Snowflake’s docs note you must list full hostnames; wildcards like *.pypi.org are not supported.) If you omit the port, it defaults to 443 for HTTPS. You could also allow all internet (0.0.0.0:443) but that’s usually too permissive.</p>
<p>Alternatively, Snowflake provides a managed rule SNOWFLAKE.EXTERNAL_ACCESS.PYPI_RULE that already includes the necessary PyPI domains. For brevity, we can use that:</p>
<pre><code class="language-sql">-- Using Snowflake-managed PyPI network rule (easiest)
CREATE OR REPLACE EXTERNAL ACCESS INTEGRATION pypi_access_integration
  ALLOWED_NETWORK_RULES = (snowflake.external_access.pypi_rule)
  ENABLED = TRUE;
</code></pre>
<p>Under this approach, Snowflake created a network rule object <code>snowflake.external_access.pypi_rule</code> that covers the needed domains. This EAI now allows egress to PyPI.</p>
<hr />
<h2><strong>Creating the External Access Integration</strong></h2>
<p>After you have your network rules, create an <strong>External Access Integration (EAI)</strong> that bundles them. For example:</p>
<pre><code class="language-sql">-- Manual approach with our own network rule
CREATE OR REPLACE EXTERNAL ACCESS INTEGRATION pypi_access_integration
  ALLOWED_NETWORK_RULES = (pypi_network)
  ENABLED = TRUE;

-- Grant usage on the integration to the app developer role
GRANT USAGE ON INTEGRATION pypi_access_integration TO ROLE my_streamlit_role;
</code></pre>
<p>The first statement makes an integration named pypi_access_integration that allows traffic to the hosts in pypi_network. We then grant USAGE on the integration to the role that owns (or deploys) the Streamlit app.</p>
<p>(If you use the Snowflake-managed rule, you could skip the CREATE NETWORK RULE and just do the first SQL block above with ALLOWED_NETWORK_RULES = (snowflake.external_access.pypi_rule).)</p>
<hr />
<h2><strong>Associating the EAI with the Streamlit App</strong></h2>
<p>Finally, tell the Streamlit object to use the integration. You can do this in either <code>CREATE STREAMLIT</code> or <code>ALTER STREAMLIT</code>. For an existing app:</p>
<pre><code class="language-sql">USE ROLE my_streamlit_role;

ALTER STREAMLIT mydb.hr_schema.hr_streamlit_app
  SET EXTERNAL_ACCESS_INTEGRATIONS = (pypi_access_integration);
</code></pre>
<p>This DDL instructs Snowflake to attach the EAI to the app. Once set, whenever you deploy or update the app, Snowflake’s container environment will allow pip-install from PyPI. Note: You can also set it at CREATE STREAMLIT time with <code>EXTERNAL_ACCESS_INTEGRATIONS=(...)</code>.</p>
<hr />
<h2><strong>Summary of Commands</strong></h2>
<p>Putting it all together, a typical sequence might be:</p>
<pre><code class="language-sql">-- 1. (Optional) Use Snowflake-managed PyPI rule
-- CREATE OR REPLACE NETWORK RULE pypi_network MODE = EGRESS TYPE = HOST_PORT VALUE_LIST=('pypi.org:443','files.pythonhosted.org:443');

-- 2. Create external access integration for PyPI
CREATE OR REPLACE EXTERNAL ACCESS INTEGRATION pypi_access_int
  ALLOWED_NETWORK_RULES = (snowflake.external_access.pypi_rule)  -- or (pypi_network)
  ENABLED = TRUE;

-- 3. Grant the integration to the app's owner role
GRANT USAGE ON INTEGRATION pypi_access_int TO ROLE my_streamlit_role;

-- 4. Attach EAI to the Streamlit app
ALTER STREAMLIT mydb.hr_schema.my_app
  SET EXTERNAL_ACCESS_INTEGRATIONS = (pypi_access_int);
</code></pre>
<p>With this in place, your Streamlit container can pip-install any public package during image build.</p>
<p><strong>For example: If you want to use Plotly, now you can simply add —</strong></p>
<h3><strong>Add Plotly to</strong> <code>requirements.txt</code></h3>
<p>Inside your project folder, ensure your <code>requirements.txt</code> includes Plotly. When the container starts, it will use the <code>pypi_access_int</code> to fetch these libraries.</p>
<pre><code class="language-plaintext"># requirements.txt
streamlit&gt;=1.53.0
plotly==5.24.0
</code></pre>
<hr />
<h2><strong>End-to-End Example Flow</strong></h2>
<p>Putting it all together, the runtime flow looks like this:</p>
<ol>
<li><p><strong>User Authentication:</strong> Rakesh accesses the Streamlit app URL (Snowflake provides a link). She is prompted to log in via Snowflake (or SSO). Snowflake issues a short-lived token for the viewer, including their default role.</p>
</li>
<li><p><strong>Streamlit App Execution:</strong> Snowflake’s compute pool routes the request to the already-running container instance of the app. The app’s code (from <a href="http://app.py"><code>app.py</code></a>) starts executing for this session.</p>
</li>
<li><p><strong>Caller’s Rights Connection:</strong> At the top of the Streamlit script, we did <code>conn = st.connection("snowflake-callers-rights")</code>. This uses the secret token to establish a Snowflake session <em>as Rakesh</em>.</p>
</li>
<li><p><strong>Query Execution:</strong> When Rakesh ask “Show me my department’s data summary”, the code runs a query via <code>conn.query(...)</code>. Snowflake sees the query coming from Rakesh’s context, checks that <code>app_owner_role</code> has CALLER SELECT on the table and Rakesh has SELECT through some application role, and then executes it. Any RLS on the table sees <code>CURRENT_ROLE()</code> = <em>Rakesh’s role — HR_ROLE.</em></p>
</li>
<li><p><strong>Return Results:</strong> The query results (already filtered by RLS) go back through the connector and are displayed in the app.</p>
</li>
<li><p><strong>Display:</strong> Streamlit renders the output (charts or text) in Rakesh’s browser. At no point did the app reveal data outside Rakesh’s authorization.</p>
</li>
</ol>
<hr />
<h2>What’s Next</h2>
<p>In this article, we covered why enterprise AI apps shift from Streamlit’s Warehouse runtime to Container runtime (SPCS) and how to configure third-party packages and restricted caller’s rights. We enabled per-user data access via caller’s rights and provided the exact SQL for network rules, EAIs, and integration with Streamlit.</p>
<p>Above all, remember that running production apps on Snowflake means designing for least privilege. The combination of Restricted Caller’s Rights and External Access Integrations gives you tight control: the app can only do what you explicitly grant. Your ongoing work will likely involve refining those grants as business needs evolve, updating and scanning dependencies, and monitoring usage (e.g. via Snowflake’s EXTERNAL_ACCESS_HISTORY or SPCS logging).</p>
<p>By following these practices, your organization can leverage AI on its Snowflake data securely and scalably, effectively turning your enterprise data warehouse into an AI-powered insights platform.</p>
<p>However, Restricted Caller Rights are only one part of building a secure enterprise AI application.</p>
<p>The next step is ensuring that every user only sees the data they are authorized to access, regardless of whether they’re interacting through a SQL query, a Cortex Agent, a Cortex Search Service, or a semantic model.</p>
<p>In the next article, we’ll take a deep dive into Row Access Policies and Dynamic Data Masking in Snowflake, where we’ll build a real-world enterprise authorization framework.</p>
<p><strong>We’ll cover topics such as:</strong></p>
<ul>
<li><p>Implementing Row Access Policies using <code>CURRENT_ROLE()</code> and <code>CURRENT_USER()</code></p>
</li>
<li><p>Restricting data by Business Unit, Department, or Country</p>
</li>
<li><p>Applying Dynamic Data Masking to sensitive fields such as salaries, invoice amounts, TCV (Total Contract Value), and confidential financial metrics</p>
</li>
<li><p>Integrating Row Access Policies with Cortex Agents and enterprise AI applications</p>
</li>
<li><p>Ensuring AI responses respect the same security policies as traditional SQL queries</p>
</li>
<li><p>Best practices for designing scalable, maintainable, and auditable security models in Snowflake</p>
</li>
</ul>
<p>By the end of the next article, you’ll have a robust authorization framework that allows AI applications to deliver personalized responses while ensuring sensitive enterprise data remains protected.</p>
<p>#snowflake #snowflakecallerrights #snowflakecortex #cortexagent #restrictedcallerrights #streamlitinsnowflake #enterpriseAI #AIagent</p>
]]></content:encoded></item><item><title><![CDATA[Securing Enterprise AI Pipelines: Using Azure Key Vault and AES Encryption for Sensitive Documents in Snowflake/Databricks AI Workloads]]></title><description><![CDATA[Organizations are rapidly building AI-powered applications on top of business documents such as contracts, invoices, employee records, financial reports, sales data, and operational data. While signif]]></description><link>https://harshtrivedii.hashnode.dev/securing-enterprise-ai-pipelines-using-azure-key-vault-and-aes-encryption-for-sensitive-documents-in-snowflake-databricks-ai-workloads</link><guid isPermaLink="true">https://harshtrivedii.hashnode.dev/securing-enterprise-ai-pipelines-using-azure-key-vault-and-aes-encryption-for-sensitive-documents-in-snowflake-databricks-ai-workloads</guid><dc:creator><![CDATA[Harsh Trivedi]]></dc:creator><pubDate>Mon, 22 Jun 2026 15:23:56 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/d606fc93-2259-47cf-8340-acb7974d6e6b.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Organizations are rapidly building AI-powered applications on top of business documents such as contracts, invoices, employee records, financial reports, sales data, and operational data. While significant effort is spent on developing AI models and intelligent agents, security is often treated as an afterthought.</p>
<p>One of the biggest risks in enterprise AI is exposing sensitive information during document storage, processing, and retrieval. Simply storing files in cloud storage or data platforms is not enough. The data should also be secured during transmission.</p>
<p>In this article, I will walk you through a security-first architecture that combines AES encryption, Azure Key Vault, Service Principals, RBAC permissions, and controlled decryption workflows to protect sensitive documents while still enabling AI-powered document extraction, analytics, and intelligent agents.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/e5e9c5f7-8065-4104-8507-fe5e10753aa5.png" alt="" style="display:block;margin:0 auto" />

<p><strong>We'll cover:</strong></p>
<p>✔ The security problem enterprises face</p>
<p>✔ Why encrypted storage alone is not enough</p>
<p>✔ Why encryption keys should never be stored in code or .env files</p>
<p>✔ Azure Key Vault fundamentals</p>
<p>✔ Service Principal authentication</p>
<p>✔ RBAC permissions and access control</p>
<p>✔ Creating and managing Key Vault secrets</p>
<p>✔ Python implementation examples</p>
<p>✔ Snowflake and Databricks integration patterns</p>
<p>✔ Secure document processing for AI workloads</p>
<p>✔ Security benefits and best practices</p>
<p>This architecture can be applied across Snowflake, Databricks, Microsoft Fabric, Azure AI Foundry, Synapse Analytics, and custom AI platforms.</p>
<hr />
<h1><strong>The Growing Security Challenge in Enterprise AI</strong></h1>
<p>Organizations are increasingly building AI applications on top of business-critical documents.</p>
<p>Examples include:</p>
<ul>
<li><p>Customer contracts</p>
</li>
<li><p>Legal agreements</p>
</li>
<li><p>Invoices</p>
</li>
<li><p>Financial statements</p>
</li>
<li><p>Employee records</p>
</li>
<li><p>Payroll information</p>
</li>
<li><p>Vendor documents</p>
</li>
<li><p>HR reports</p>
</li>
<li><p>Employee attrition predictions</p>
</li>
<li><p>Compliance reports</p>
</li>
<li><p>Healthcare records</p>
</li>
<li><p>Procurement documents</p>
</li>
</ul>
<p>These documents often contain highly sensitive information.</p>
<p>At the same time, organizations want AI systems to:</p>
<ul>
<li><p>Extract information from documents</p>
</li>
<li><p>Build search experiences</p>
</li>
<li><p>Power RAG solutions</p>
</li>
<li><p>Support AI agents</p>
</li>
<li><p>Generate business insights</p>
</li>
<li><p>Answer natural language questions</p>
</li>
<li><p>Generate KPIs, charts and visuals</p>
</li>
</ul>
<p>This creates a difficult challenge:</p>
<blockquote>
<p>How can we allow AI systems to use sensitive information without exposing the underlying documents?</p>
</blockquote>
<hr />
<h2><strong>The Wrong Approach</strong></h2>
<p>Many organizations unknowingly introduce security risks by:</p>
<ul>
<li><p>Storing documents unencrypted</p>
</li>
<li><p>Storing encryption keys in source code or environment files</p>
</li>
<li><p>Keeping secrets in configuration files</p>
</li>
<li><p>Saving credentials inside notebooks</p>
</li>
<li><p>Embedding keys inside ETL pipelines</p>
</li>
</ul>
<p>For example:</p>
<pre><code class="language-plaintext">AES_KEY = "MySecretEncryptionKey"
</code></pre>
<p>If someone gains access to:</p>
<ul>
<li><p>GitHub repositories</p>
</li>
<li><p>DevOps pipelines</p>
</li>
<li><p>Virtual machines</p>
</li>
<li><p>Databases</p>
</li>
<li><p>Storage accounts</p>
</li>
</ul>
<p>they may gain access to both the encrypted documents and the encryption key.</p>
<p>At that point, encryption provides little value.</p>
<hr />
<h2><strong>The Security Principle</strong></h2>
<p>One of the most important security principles is:</p>
<blockquote>
<p>Encrypt the data and protect the encryption key separately.</p>
</blockquote>
<p>Even if attackers gain access to encrypted documents, they should not be able to access the encryption key.</p>
<p>This is where Azure Key Vault becomes critical.</p>
<hr />
<h1><strong>The Solution Architecture</strong></h1>
<p>The architecture consists of five major layers:</p>
<ol>
<li><p>Document Encryption</p>
</li>
<li><p>Secure Key Storage</p>
</li>
<li><p>Controlled Authentication</p>
</li>
<li><p>In-Memory Decryption</p>
</li>
<li><p>AI Processing</p>
</li>
</ol>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/66ccaa55-c005-42f1-9d19-51c56c6688a4.png" alt="" style="display:block;margin:0 auto" />

<p><strong>Document Source -</strong> The source of document can be anything like SAP, Cloud storage, Azure Blob Storage etc.</p>
<p>When we move them to our storage layer, they should be encrypted.</p>
<h1><strong>A) Why We Encrypt Documents</strong></h1>
<p>Consider documents such as:</p>
<ul>
<li><p>Client contracts</p>
</li>
<li><p>Employee compensation data</p>
</li>
<li><p>Vendor agreements</p>
</li>
<li><p>Invoices</p>
</li>
<li><p>Strategic business reports</p>
</li>
<li><p>Attrition prediction outputs</p>
</li>
</ul>
<p>These files often contain information that should not be accessible to every user or system.</p>
<p>Before storing these files, we encrypt them using AES-256.</p>
<p>This ensures that even if storage is compromised or a developer has RBAC (Role based access control), the document remains unreadable without the encryption key.</p>
<hr />
<h1><strong>B) Creating an AES Encryption Key</strong></h1>
<p>Example:</p>
<pre><code class="language-python">from Crypto.Random import get_random_bytes

AES-256 = 32 bytes

aes_key = get_random_bytes(32)

print(aes_key.hex()) 32 bytes = 256 bits
</code></pre>
<p>This creates an AES-256 encryption key.</p>
<h3><strong>Why AES?</strong></h3>
<p>AES (Advanced Encryption Standard) is one of the most widely used symmetric encryption algorithms in the world.</p>
<p>Advantages:</p>
<ul>
<li><p>Fast</p>
</li>
<li><p>Highly secure</p>
</li>
<li><p>Industry standard</p>
</li>
<li><p>Suitable for large documents</p>
</li>
<li><p>Supported across cloud platforms</p>
</li>
</ul>
<p>AES-256 is commonly used in enterprise environments.</p>
<hr />
<h1>C) What is Azure Key Vault?</h1>
<p>Azure Key Vault is a managed Azure service designed for securely storing and controlling access to secrets, cryptographic keys, and certificates.</p>
<p>Instead of storing sensitive values in application code, organizations place them inside a centralized vault protected by Azure Identity and RBAC controls.</p>
<p>Azure Key Vault supports three primary object types:</p>
<p><strong>Secrets</strong></p>
<p><strong>Examples:</strong></p>
<ol>
<li><p>API Keys</p>
</li>
<li><p>Database Passwords</p>
</li>
<li><p>OAuth Client Secrets</p>
</li>
<li><p>AES Encryption Keys</p>
</li>
</ol>
<p><strong>Keys</strong></p>
<p><strong>Examples:</strong></p>
<ol>
<li><p>RSA Keys</p>
</li>
<li><p>Cryptographic Signing Keys</p>
</li>
<li><p>Customer Managed Encryption Keys</p>
</li>
</ol>
<p><strong>Certificates</strong></p>
<p><strong>Examples:</strong></p>
<ol>
<li><p>SSL Certificates</p>
</li>
<li><p>TLS Certificates</p>
</li>
</ol>
<p>In our use case, the AES encryption key is stored as a Secret.</p>
<p><strong>Why Azure Key Vault is required?</strong></p>
<ol>
<li><p>Centralized secret management</p>
</li>
<li><p>Fine-grained access control</p>
</li>
<li><p>Secret rotation support</p>
</li>
<li><p>Access auditing</p>
</li>
<li><p>Compliance support</p>
</li>
<li><p>Reduced insider risk</p>
</li>
</ol>
<h3>Creating Azure Key Vault</h3>
<p><strong>1.Navigate to:</strong> portal.azure.com</p>
<blockquote>
<p>Azure Portal -&gt; Create Resource -&gt; Key Vault -&gt; Create</p>
</blockquote>
<p><strong>2.Configure:</strong></p>
<p><code>Subscription</code></p>
<p><code>Resource Group</code></p>
<p><code>Region (eg. Central India)</code></p>
<p><code>Vault Name (eg. projectname-aes-key-kv)</code></p>
<p><strong>3.For Authorization/Permission Model:</strong></p>
<p><code>Select: Azure Role Based Access Control (RBAC)</code></p>
<p><strong>4.For Networking:</strong></p>
<p><code>Public access: Selected networks</code></p>
<p>or</p>
<p><code>Private endpoint</code></p>
<p><strong>5.Permissions Required to Create a Key Vault</strong></p>
<p>Typically one of the following roles:</p>
<ol>
<li><p>Owner</p>
</li>
<li><p>Contributor</p>
</li>
<li><p>Key Vault Contributor</p>
</li>
</ol>
<p><strong>Important:</strong></p>
<p>Key Vault Contributor can manage the Key Vault resource itself but cannot read secrets stored within the vault.</p>
<p>This separation improves security.</p>
<p><strong>6.Storing the AES Key in key vault secret and permission required to create secrets.</strong></p>
<p>IMP: When using Azure RBAC, you will need Key Vault secret officer role to create secret.</p>
<p>Once the role is assigned, Create and Store the generated AES key as a secret.</p>
<p><strong>Example:</strong></p>
<p>Secret Name:</p>
<p><code>AES-DOCUMENT-ENCRYPTION-KEY : secret _code</code></p>
<p>The key is now centrally managed and protected.</p>
<hr />
<h1>D) Creating a Service Principal Applications, requires an identity to access Azure resources.</h1>
<p>This is where Service Principals are used.</p>
<p>Think of a Service Principal as:</p>
<blockquote>
<p>A non-human identity used by applications.</p>
</blockquote>
<p><strong>Example:</strong></p>
<p><mark class="bg-yellow-200 dark:bg-yellow-500/30">Snowflake/Databricks/Azure Function/Fabric Pipeline/Custom API</mark></p>
<p><mark class="bg-yellow-200 dark:bg-yellow-500/30">↓</mark></p>
<p><mark class="bg-yellow-200 dark:bg-yellow-500/30">Service Principal</mark></p>
<p><mark class="bg-yellow-200 dark:bg-yellow-500/30">↓</mark></p>
<p><mark class="bg-yellow-200 dark:bg-yellow-500/30">Azure Key Vault</mark></p>
<p><mark class="bg-yellow-200 dark:bg-yellow-500/30">↓</mark></p>
<p><mark class="bg-yellow-200 dark:bg-yellow-500/30">AES Secret</mark></p>
<p>The Service Principal becomes the trusted identity that accesses the vault.</p>
<p><strong>Why Service Principals Are Important</strong></p>
<p><strong>Without Service Principals:</strong></p>
<ol>
<li><p>Applications would require user credentials</p>
</li>
<li><p>Automation becomes difficult</p>
</li>
<li><p>Auditing becomes harder</p>
</li>
</ol>
<p><strong>Service Principals provide:</strong></p>
<ol>
<li><p>Automated authentication</p>
</li>
<li><p>Secure access</p>
</li>
<li><p>RBAC integration</p>
</li>
<li><p>Least privilege access</p>
</li>
</ol>
<p><strong>RBAC access for service principal on the key vault: Application Access</strong></p>
<p>For Snowflake, Databricks, Fabric, or AI platforms:</p>
<p>Role:</p>
<p><code>Key Vault Secrets User</code></p>
<p>Can:</p>
<ol>
<li><p>Read secret values</p>
</li>
<li><p>Read secret metadata</p>
</li>
</ol>
<hr />
<h1>E) Python Example: Reading the AES Key</h1>
<pre><code class="language-python">from azure.identity import ClientSecretCredential
from azure.keyvault.secrets import SecretClient

credential = ClientSecretCredential(
    tenant_id="&lt;tenant-id&gt;",
    client_id="&lt;client-id&gt;",
    client_secret="&lt;client-secret&gt;"
)

vault_url = "https://company-kv.vault.azure.net"

client = SecretClient(
    vault_url=vault_url,
    credential=credential
)

aes_key = client.get_secret(
    "AES-DOCUMENT-ENCRYPTION-KEY"
).value
</code></pre>
<p>The application retrieves the key only when needed.</p>
<hr />
<h1>F) Encrypting Sensitive Documents</h1>
<pre><code class="language-python">from Crypto.Cipher import AES
from Crypto.Util.Padding import pad
from Crypto.Random import get_random_bytes

# AES key retrieved from Azure Key Vault
key = bytes.fromhex(aes_key_hex)

# Generate random IV (16 bytes for AES)
iv = get_random_bytes(16)

with open("contract.pdf", "rb") as file:
    plaintext = file.read()

cipher = AES.new(key, AES.MODE_CBC, iv)

ciphertext = cipher.encrypt(
    pad(plaintext, AES.block_size)
)

# Store IV along with encrypted content
encrypted_content = {
    "iv": iv,
    "ciphertext": ciphertext
}
</code></pre>
<p>The encrypted file can now be stored safely.</p>
<hr />
<h1>G) AI Processing Pattern</h1>
<h3>When processing begins:</h3>
<ol>
<li><p>Retrieve AES key from Key Vault</p>
</li>
<li><p>Decrypt document in memory</p>
</li>
<li><p>Extract structured or unstructured chunk information</p>
</li>
<li><p>Persist extracted results</p>
</li>
<li><p>Discard decrypted content</p>
</li>
</ol>
<p><strong>The key principle is:</strong></p>
<blockquote>
<p>Decrypt only when required and only in memory.</p>
</blockquote>
<hr />
<h1>H) Platform Integration Examples</h1>
<img src="https://cdn.hashnode.com/uploads/covers/6a39322015ed6345c6f1c5cc/ca3c84d4-da73-434f-8927-d6e4c7694353.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h1>I) AI Agents and RAG Systems</h1>
<p>Once document information is extracted into structured datasets, AI systems can answer questions such as:</p>
<ol>
<li><p>Which contracts expire next month?</p>
</li>
<li><p>Which invoices exceed $100,000?</p>
</li>
<li><p>Which vendors require renewal?</p>
</li>
<li><p>What is the predicted attrition risk for the Sales department?</p>
</li>
<li><p>Which customers have pending obligations?</p>
</li>
</ol>
<p>The AI model operates on structured business data rather than unrestricted access to sensitive raw documents.</p>
<hr />
<h1>J) Security Benefits Achieved</h1>
<h3>This architecture provides:</h3>
<ol>
<li><p><strong>Encryption at Rest -</strong> Documents remain protected in storage.</p>
</li>
<li><p><strong>Key Separation -</strong> Encryption keys are separated from encrypted data.</p>
</li>
<li><p><strong>Least Privilege Access -</strong> Applications receive only required permissions.</p>
</li>
<li><p><strong>Auditing -</strong> Secret access can be monitored.</p>
</li>
<li><p><strong>Key Rotation -</strong> Encryption keys can be replaced centrally.</p>
</li>
<li><p><strong>Compliance Support -</strong> Supports enterprise governance and security requirements.</p>
</li>
<li><p><strong>Reduced Attack Surface -</strong> Compromised storage does not automatically expose sensitive information.</p>
</li>
</ol>
<hr />
<h1>Final Thoughts</h1>
<p>As organizations accelerate AI adoption, securing the data feeding those AI systems becomes just as important as building the models themselves.</p>
<p>A secure AI platform is not only about model accuracy, vector databases, or intelligent agents. It is equally about protecting contracts, invoices, employee information, and other sensitive business assets that power those systems.</p>
<p>By combining AES encryption, Azure Key Vault, Service Principals, RBAC permissions, and controlled in-memory decryption, organizations can build AI platforms that are secure, scalable, compliant, and enterprise-ready.</p>
<p><strong>The goal is simple:</strong></p>
<blockquote>
<p>Protect the documents.</p>
</blockquote>
<blockquote>
<p>Protect the keys.</p>
</blockquote>
<blockquote>
<p>Allow AI access only when necessary.</p>
</blockquote>
<hr />
<h2>What next?</h2>
<p>In my next article, I will walk you through how we implemented enterprise-grade access controls for AI-powered platforms using techniques such as:</p>
<ul>
<li><p>Row-Level Security (RLS)</p>
</li>
<li><p>Row Access Policies</p>
</li>
<li><p>Dynamic Data Masking</p>
</li>
<li><p>Role-Based Access Control (RBAC)</p>
</li>
<li><p>Business Unit-based filtering</p>
</li>
<li><p>Customer-based entitlements</p>
</li>
<li><p>Financial data masking</p>
</li>
<li><p>Secure document access controls</p>
</li>
<li><p>Time-limited document download links</p>
</li>
<li><p>Agent-level authorization patterns</p>
</li>
<li><p>Governance frameworks for AI applications</p>
</li>
</ul>
<p>We'll explore how an AI agent can return different answers to different users for the same question, while ensuring that sensitive information remains protected and compliant with organizational policies.</p>
<p>Because securing the documents is important.</p>
<p>But securing who can access the information extracted from those documents is what truly makes an AI platform enterprise-ready.</p>
<p>#ArtificialIntelligence #DataEngineering #Azure #AzureKeyVault #CyberSecurity #DataSecurity #DocumentAI #Snowflake #Databricks #EnterpriseAI #AIAgents #AI</p>
]]></content:encoded></item></channel></rss>