Data Warehouse Integration Technical Process - Databricks

Last updated: September 29, 2026

Overview

CRED's Databricks integration syncs tables from your Databricks workspace into CRED as Accounts, Contacts, Companies, People, Deals, or Leads, on SOC II-compliant infrastructure.

Databricks connects through Polytomic, CRED's managed data-integration partner. This is different from the native CRM integrations (Salesforce, HubSpot, Dynamics 365), which call the CRM's API directly:

  • You choose the tables. Databricks has no fixed objects, so you pick each table to sync and the CRED record type it becomes.

  • Field discovery is automatic. CRED reads the columns from each table's schema.

  • Syncs run continuously. Polytomic watches the selected tables and brings in new and changed rows as it finds them.

  • Each workspace is isolated. Every CRED workspace gets its own Polytomic organisation, credentials, and storage area.

How long setup takes depends on table size and SQL warehouse capacity. Large tables can take several hours to finish their first sync.

How It Works

Solid arrows show the read path. The dashed line is the write-back path, which is still to be confirmed.

  1. Polytomic reads the tables you selected in Databricks and copies them into a staging area that belongs only to your CRED workspace.

  2. CRED reads the staging area. After the first full sync, it uses each table's tracking field (e.g. updated_at) to pick up only new or changed rows.

  3. The rows become CRED records of the entity type you chose. CRED then matches them to its enriched Company and Person data.

Supported Entities

Each Databricks table maps to one CRED entity type. To bring in several record types, create one sync per table.

CRED Entity

Databricks Source

Import

Export

Company

Any table or view you select

Yes

[VERIFY]

Person

Any table or view you select

Yes

[VERIFY]

Account

Any table or view you select

Yes

[VERIFY]

Contact

Any table or view you select

Yes

[VERIFY]

Deal

Any table or view you select

Yes

[VERIFY]

Lead

Any table or view you select

Yes

[VERIFY]

Table Requirements

Every table you sync needs these three columns:

Requirement

Why

Notes

Primary key: one unique column

Identifies each record so CRED can update it instead of duplicating it

CRED detects a declared primary key automatically, or a column named id or uuid. You can also pick one yourself. Composite (multi-column) primary keys are not supported. For those tables, build a view with a single surrogate key.

Tracking field: a date or timestamp column

Lets CRED sync only new or changed rows after the first run

Must be updated whenever a row changes, e.g. updated_at or last_modified.

Record name field

The display name shown for each record in CRED

e.g. company_name, full_name.

Prerequisites

On the Databricks side

Have a Databricks admin on hand for the first part of setup.

  • Workspace hostname, e.g. <workspace>.cloud.databricks.com (AWS/GCP) or adb-<id>.<n>.azuredatabricks.net (Azure)

  • SQL warehouse HTTP path (serverless, pro, or classic), from SQL Warehouses → [warehouse] → Connection details

  • Access token for a service principal or shared service account. Its read permission (USE CATALOG, USE SCHEMA, SELECT) must cover every catalog, schema, and table you want to sync

  • [VERIFY] If write-back is on, the same identity needs MODIFY on the target tables

  • Network access: if the workspace uses an IP access list or Private Link, allowlist the CRED/Polytomic egress IPs [VERIFY: add IPs]

Recommendation: Use a service principal or shared service account rather than a personal user token. A token tied to one person stops working when they leave or their access changes.

On the CRED side

  • A CRED user with permission to manage Workspace Integrations. If you're not sure your role has it, ask your CRED account owner.

How to Connect Databricks

Step 1: Access Workspace Integrations

Go to My Account → Workspace Integrations, or visit commercial.credplatform.com/my-account/workspace-integrations. Scroll to Data Warehouses & Analytics Platforms, select the Databricks tile, and click Connect.

Step 2: Enter Your Databricks Connection Details

A connection window opens as a popup. It is hosted by Polytomic, CRED's integration partner. If nothing appears, allow popups for credplatform.com and try again.

Enter:

  • Server hostname: your workspace hostname, without https://

  • HTTP path: the SQL warehouse HTTP path

  • Access token: the service principal or service-account token

Submit the form. The connection is checked, the popup closes, and you'll see a "Data source connected successfully" notification. The Databricks tile now shows as connected.

Your access token is stored encrypted in the integration partner's secure credential store. It is used only by the sync engine and CRED staff can't see it.

Step 3: Create a Data Sync

Go to My Account → Data Sync (commercial.credplatform.com/my-account/data-sync). Your Databricks connection appears under Data Sources with status Active. Click New Sync to open the Configure Data Sync wizard.

Step 4: Configure with the Data Sync Wizard

The wizard has four steps:

  1. Select Source: choose your Databricks connection, then the schema or table to sync.

  2. Select Fields: choose the columns to bring into CRED. Include the tracking field, the record name field, and the primary key if you're setting it yourself.

  3. Configure Mapping:

    • Entity Type: the kind of CRED record these rows are (Company, Person, Contact, Account, Deal, Lead)

    • Record Name Field (required): the column used as each record's display name

    • Tracking Field (required): a date or timestamp column used for incremental sync

    • Primary Key Field (optional): leave on Auto-detect unless the table has no declared key and no id or uuid column

  4. Review & Create: check the summary and click Create.

Repeat for each table you want to sync.

Step 5: Initial Data Sync

The first full sync of the table starts automatically once the sync is created. You can follow progress on the Data Sync page, where available tables and row counts appear as each sync finishes. After that, only changed rows sync.

Sync Behaviour

Sync Schedule

Direction

Frequency

Trigger

Databricks → CRED staging

Continuous

Automatic (Polytomic bulk sync)

CRED staging → CRED records

Incremental, on CRED's data-sync schedule [VERIFY: interval]

Automatic

CRED → Databricks (write-back)

Continuous [VERIFY]

Automatic (Polytomic model sync)

Timings may change. Continuous syncs keep your SQL warehouse busy, so check its auto-stop and sizing settings. For a custom schedule, contact your CRED representative.

How Data Flows

  • Each Databricks row syncs 1:1 to a CRED record of the chosen entity type, identified by its primary key

  • Account and Company rows are matched to CRED Companies with enriched data. Contact and Person rows are matched to CRED Persons [VERIFY: same matching pipeline as CRM imports]

  • Synced accounts appear under Company → My Accounts, and synced contacts under People → My Contacts

  • Updated rows (with a newer tracking-field value) update the existing CRED record. They don't create duplicates

  • Schema changes such as new columns are picked up the next time fields are read. To bring a new column into CRED, add it to the sync [VERIFY: whether an existing sync can be edited, or must be recreated]

Field Settings

Data Type Mapping

Databricks Type

CRED Field Type

STRING, VARCHAR, CHAR, BINARY

Text

TINYINT, SMALLINT, INT, BIGINT

Number

FLOAT, DOUBLE, DECIMAL

Number

BOOLEAN

Checkbox

DATE

Date

TIMESTAMP, TIMESTAMP_NTZ

Timestamp

STRUCT, MAP

Record

ARRAY

Multi-value [VERIFY]

VARIANT / JSON strings

Text

Types are normalised on the way through the staging area. Map other types, such as geospatial columns, to Text in a view before syncing.

Special Fields

Field

Purpose

Primary key

Unique record ID. Used to match rows on update and on write-back

Tracking field

Incremental sync. CRED also sets this to the current time when it writes a change back [VERIFY]

Record name

The display name in CRED lists and sidebars

Required Fields

  1. Required in Databricks: columns defined NOT NULL. A write-back fails if CRED doesn't have a value for them.

  2. Required in CRED: the primary key and record name must have values. Rows with a null primary key are skipped and reported as sync errors.

Additional Field Controls

  • Is Exportable?: an admin setting that controls whether a field's data appears in CSV exports

  • Column selection: only the columns you pick in Select Fields are synced. To keep sensitive columns (PII, financials) out of CRED, leave them unselected, or sync from a view that leaves them out

  • Sync Limits: depend on your CRED plan. Contact your representative for your limits

Matching & Deduplication

  • CRED matches imported Accounts and Companies to its own Company data by domain and other identifiers. Include a website or domain column to improve match rates

  • Matching statistics appear under My Account → Settings → Matching

  • Click a row to change which CRED Company an imported record matches

  • Monitor unmatched records and find their matches by hand

  • CRED flags possible duplicates for review. It doesn't merge them automatically

  • Inside a single table, the primary key prevents duplicates. If two tables are mapped to the same entity type, their records aren't linked automatically

Go-Live Checklist

Before rolling the integration out to your team:

  • The service principal or token has USE CATALOG, USE SCHEMA, and SELECT on every table to sync (and MODIFY if write-back is on [VERIFY])

  • The SQL warehouse is sized for continuous syncs, and its auto-stop setting has been reviewed

  • IP allowlist / Private Link rules include the CRED/Polytomic egress IPs

  • Every table has one unique primary key, a reliable tracking field, and a record name field

  • Each table is mapped to the right CRED entity type

  • Sensitive columns are excluded, or left out through a view

  • The first sync finished and row counts on the Data Sync page match Databricks

  • Sample records have been checked in CRED, including match to the right Company or Person

  • Write-back tested with a sample record in a non-production table [VERIFY]

  • Sales and support teams trained on where synced data appears, matching, and troubleshooting

  • Is Exportable? settings reviewed for sensitive fields

Disconnecting Databricks

  1. Go to My Account → Data Sync

  2. Find the Databricks connection under Data Sources and click Remove

  3. Confirm in the Remove Data Source dialog

You can also open My Account → Workspace Integrations, find the Databricks tile, and use its settings.

Disconnecting stops all syncing in both directions. Records already in CRED stay but no longer update. Your Databricks data isn't changed. To reconnect, you'll need to enter credentials again and wait for a fresh initial sync.

Tip: To stop CRED's access at the Databricks end as well, revoke the access token or remove the service principal's grants in Databricks.

Troubleshooting

Connection Status

Status

Meaning

Active

Connected and syncing

Inactive

Connected but not syncing

Relink Needed

Credentials were rejected, e.g. the token expired or was revoked. Reconnect from Workspace Integrations

Common Issues

Symptom

Likely cause

What to check

Connect popup doesn't open

The browser is blocking popups

Allow popups for credplatform.com and click the tile again

Popup shows an authentication error

Wrong hostname or HTTP path, or an invalid or expired token

Check the hostname (no https://) and HTTP path against Connection details. Generate a new token if needed

Connected, but no schemas or tables in the wizard

The token can't see the catalog or schema

Run SHOW GRANTS ON SCHEMA <catalog>.<schema> and grant USE CATALOG / USE SCHEMA / SELECT to the principal

"Composite primary key not supported"

The table's declared primary key has more than one column

Sync from a view with a single surrogate key, e.g. concat_ws('-', col_a, col_b) AS id

Wizard won't continue past Configure Mapping

No record name or tracking field chosen, or it wasn't picked in Select Fields

Go back to Select Fields and include those columns

Records not updating after the first sync

Tracking field not updated when rows change

Make sure your pipeline sets updated_at (or equivalent) on every change

Some rows missing

Null primary key values, or row-level security hiding rows from the principal

Check for null keys, and review row filters and column masks in Unity Catalog

Sync slow, stalled, or incomplete

The warehouse auto-stopped or is too small, or queries are queued

Make the warehouse bigger, lengthen auto-stop, or use serverless. Contact CRED support to review sync logs

"Permission denied" on specific tables

Unity Catalog grants, row filters, or column masks

Grant the required privileges, or take the table out of the sync

Filing a Support Ticket

Include the following when you report a sync issue:

  • Databricks table name (catalog.schema.table) and the CRED entity type it's mapped to

  • The primary key, tracking field, and record name field you chose

  • Screenshot of the affected record in CRED, and the matching row in Databricks

  • What you expected to happen and what actually happened

  • Steps to reproduce

  • Status shown on the Data Sync page, and any error message

Important Notes

  • This is a living reference and will be updated as the integration changes

  • The Databricks connection runs through Polytomic, CRED's managed integration partner. Each CRED workspace gets its own Polytomic organisation and credentials

  • Synced data is stored in a segregated, encrypted instance that only permitted organisations can access

  • Data residency depends on the workspace. Confirm your region with your CRED representative. If you have EU-only or in-region requirements, ask for the current data-processing addendum before connecting

  • Unity Catalog permissions, row filters, and column masks still apply, because CRED reads with the credentials you provide

  • CRED doesn't query Databricks live when you use the app. Data refreshes through the sync

  • Sync timings are approximate and can change

  • For custom configurations, schedule changes, or enabling features, contact your CRED representative