DataLens how-to guides
23 step-by-step guides covering the whole path from a raw file or database to a governed, published dataset. Every step describes a screen that exists in the product today.
Where to start
If you are new, work through them in order: load a dataset, clean it, model it, then govern and publish it. That sequence is the product's own lifecycle — Data, Prepare, Model, Orchestrate, Analyse, Deliver, Govern — and each guide picks up where the previous one stopped.
If you already have data in DataLens, go straight to the guide for the task in front of you. Each one lists its own prerequisites, so you will know immediately whether you need to do something else first.
Guides marked Pro describe capabilities on the paid tier. They are documented in full so you can see what is behind them before deciding whether you need it.
Get data in
Uploads, databases, APIs, files, change capture, CRM, lakes, documents and AI providers.
Upload a CSV or Excel file
Load a spreadsheet or delimited file and get a profiled dataset you can clean, model and chart.
Connect a database
Connect PostgreSQL, MySQL, SQL Server, Redshift, BigQuery or Snowflake, browse the tables and load what you need.
Pull data from a REST API
Call a JSON API with the auth it needs, point at the array inside the response, and load the result as a dataset.
Load files from FTP or SFTP
Connect an FTP or SFTP server, browse directories, preview a file and load it as a dataset.
Set up CDC sync
Keep a dataset current by capturing changes from a source table on a schedule, rather than reloading it whole.
Connect a CRM
Bring Salesforce, HubSpot, Dynamics 365 or Zoho objects in as datasets you can join with everything else.
Connect a data lake
Read from S3, ADLS Gen2, Google Cloud Storage, Databricks, Dremio, Trino or MinIO — your storage, your control.
Build a document collection
Load Word, Markdown, HTML or text documents as a collection — stored whole, not parsed into rows.
Connect an AI model
Add a model provider with your own credentials, choose which models are allowed, and set a default.
Prepare, model and orchestrate
Cleaning, transforms, joins, source and target models, mappings and pipelines.
Clean a messy dataset
Drop nulls, fill gaps, remove duplicates and cap outliers — with a live preview before anything is applied.
Transform and derive columns
Cast types, split and rename columns, standardise names and values, and build new columns from expressions.
Join datasets and create views
Join datasets on the keys that relate them and save the result as a reusable view.
Explore a source model
See the tables, columns and relationships of a system you did not design — as a diagram you can read.
Map source to target
Connect source columns to target columns, see what is unmapped, and export a source-to-target mapping.
Build a data pipeline
Turn a sequence of steps into a repeatable pipeline with a defined execution order.
Govern, analyse and deliver
Profiling, PII, quality rules, lineage, charts, simulation, documents, exports and data products.
Profile a dataset and find PII
See what is really in each column, and find the personal data before it ends up somewhere it should not.
Set data quality rulesPro
Turn a one-off quality check into a standing rule, so the same problem is caught every time.
Trace where a column came from
Follow a field back from the dataset you are looking at to the source it came from, step by step.
Chart and pivot a dataset
Bar, line and scatter charts, histograms, a correlation matrix and a pivot table over your data.
Run a scenario simulation
Ask what happens if a variable changes — forecasts, shocks, comparisons and impact, on your own data.
Publish and export data
Export as CSV, TSV, JSON, JSONL or Excel, publish an API endpoint, or generate a PDF or Word report.
Generate a data document
Produce an architecture document, a model spec, a mapping or a runbook from what you have already built.
Build a data productPro
Turn scattered datasets into a defined, governed product with a named owner and a documented schema.
Follow these on your own data
DataLens is in private beta. Request access and every guide here becomes something you can do rather than read.
Request beta access