> For the complete documentation index, see [llms.txt](https://dots.gitbook.io/dots-docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://dots.gitbook.io/dots-docs/project-prep/data-connectors.md).

# Data Connectors

How you connect your datasets is one of the first decisions you will make on the platform, and it shapes the questions you will be able to answer later.

### Why do data connectors matter?

Most organisations do not work with a single dataset. You are likely holding some combination of surveys, interviews, observation notes, field reports, monitoring data, support tickets, and CRM records, often collected by different teams at different points in time. Each one tells you something on its own, but the questions worth asking usually sit in the space between them.

A data connector describes how those datasets relate to each other once they are on the platform. The connector you choose decides whether you can compare across studies, roll findings up to a single unit, or follow the same person or place from one stage to the next. There is no single correct answer here. The right connector is the one that matches how your data was collected and what you need to get out of it.

Let's understand with an example:

Your water sanitation and hygiene (WASH) research organisation is conducting a cross-country comparative study (India & Kenya) on household sanitation practices. Across the study you collect household surveys, in-depth interviews, community observation notes, and monthly facility monitoring records.

You could keep each of these entirely separate, so that the survey analysis and the interview analysis never meet. You could tag them all by district and programme so that you can compare them side by side. You could anchor them to a household register so that every piece of data rolls up to a household. Or you could chain them in the order they were collected, so that you can trace a household from first contact through to reported outcomes.

Each of those choices is a data connector.

***What are the different data connector styles possible on the Dots platform?***

There are five. They differ in how tightly your datasets are bound to one another, from no connection at all through to a full sequence.

### 1. Isolated datasets

With this connector each dataset lives on its own. There are no links between them, no shared structure, and no roll-up analysis across them.

<figure><img src="https://2208645187-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FgczwnwqbMALFEXVtg0t1%2Fuploads%2F1nryJIaijE0USz8Yw6v2%2Fimage.png?alt=media&amp;token=079395f3-f737-494d-8ede-4a58652a2639" alt=""><figcaption></figcaption></figure>

This suits situations where your datasets genuinely have nothing to do with each other: separate studies, unrelated pilots, data owned by different teams, or older projects you want to preserve exactly as they were.

What you gain is simplicity. Each project stays clean and self-contained, reporting stays independent, and there is nothing to maintain between datasets. What you give up is any view across them. You cannot compare projects, produce shared reporting, or apply one standard set of themes or metrics to everything at once.

In the WASH example, this is what you would have if the India study and the Kenya study were run by different teams and reported separately, with no intention of comparing the two.

Elsewhere this connector shows up as independent baseline studies across countries in MEL work, separate evaluations commissioned by different funders in research, or customer research projects run by different product teams.

### 2. Floating connections (shared tags)

Here your datasets stay separate, but they share common dimensions such as geography, gender, programme, product area, customer segment, or themes. Nothing connects at the record level. The datasets meet through the tags they have in common.

<figure><img src="https://2208645187-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FgczwnwqbMALFEXVtg0t1%2Fuploads%2Fp1Pr1eg5YU3i9aFhlosN%2Fimage.png?alt=media&amp;token=ecc3adbb-3fc6-4347-81b9-844c67bbb7fa" alt=""><figcaption></figcaption></figure>

This is the connector to reach for when you want comparison across sources without forcing very different datasets into a single structure. It is usually the first step organisations take toward connected analysis, because it asks very little of the data: you only need to agree on a shared vocabulary.

With shared tags in place you can filter several datasets by the same dimension, follow a theme across multiple sources, build dashboards that draw on more than one study, and set qualitative and quantitative findings next to each other. What stays out of reach is anything that depends on individual records lining up. You cannot link a specific survey response to a specific interview, or follow one household across datasets.

In the WASH example, tagging your surveys, interviews, and observation notes by district, country, and thematic framework lets you ask what people say about shared toilet access in Kenya against what they say in India, even though the three datasets are structured nothing alike.

Elsewhere: surveys, interviews and case studies all tagged by district and programme in MEL work; multiple studies running on the same thematic framework in research; app reviews, support tickets and interview transcripts sharing one set of themes in product.

### 3. Hub-and-spoke

With this connector one dataset acts as a central register, the source of truth, and every other dataset connects back to it.

<figure><img src="https://2208645187-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FgczwnwqbMALFEXVtg0t1%2Fuploads%2F3BGb4UCj3DNdaufzBYLn%2Fimage.png?alt=media&amp;token=cdfc9dfd-00fb-45f5-9eee-d5463ef81f73" alt=""><figcaption></figcaption></figure>

Reach for this when there is a common entity running through all of your data: beneficiaries, schools, health facilities, volunteers, customers, or accounts. It is the most common choice for long-running programmes, where the same units are visited and measured repeatedly over years.

Because everything points at the register, you can roll findings up to a single unit, apply consistent filters across every dataset, share metadata rather than re-entering it, and cut down on duplication. Governance gets easier too, since there is one agreed list of who or what you are studying. The cost is maintenance. The register has to be kept accurate, and changes to its structure ripple out to everything connected to it.

In the WASH example, a household register sits at the centre. Survey responses, interview transcripts, and observation notes all attach to a household, so you can open a single household and see everything you know about it, or roll every source up to district level in one view.

Elsewhere: a school register connected to surveys, observations and interviews in MEL work; a facility register connected to monitoring and feedback systems in public health; a customer account connected to support tickets, interviews and usage data in product.

### 4. Daisy chain

Here your datasets connect in sequence, with each one building on the one before it.

<figure><img src="https://2208645187-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FgczwnwqbMALFEXVtg0t1%2Fuploads%2FwKwfhkb6OCVnlzmR6bUR%2Fimage.png?alt=media&amp;token=f4240153-8b8f-4b63-98a2-75e0fd955278" alt=""><figcaption></figcaption></figure>

This fits data that follows a process, a workflow, or a journey, where the order in which things happened is part of what you are studying. It is most useful when understanding progression matters more than understanding any single dataset on its own.

A chained setup lets you trace outcomes across stages, follow how something changes over time, and connect cause to effect across datasets rather than within one. The difficulty is dependency. Each dataset relies on the one before it, so missing links break the chain and gaps in the middle can leave you unable to follow a case through to the end.

In the WASH example, this is how you would set things up to follow households from a baseline survey, through a sanitation intervention, to a follow-up interview and an endline measurement, so that you can see which households changed practice and read what they said about why.

Elsewhere: outreach to participation to outcomes to impact in MEL work; screening to interviews to follow-up interviews in research; acquisition, activation, engagement and retention in product.

### 5. Combination data connectors

In practice most organisations use more than one connector at the same time, and that is expected. They are not exclusive, and a workspace usually settles into a mix.

A **MEL setup** often runs hub-and-spoke, with surveys, interviews and observations all connected to a programme register, and floating connections layered on top so that shared themes and geographies work across everything.

<figure><img src="https://2208645187-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FgczwnwqbMALFEXVtg0t1%2Fuploads%2F4JLvOHIdidcs2zIus5wr%2Fimage.png?alt=media&amp;token=f2a53864-bb52-4bc7-ad56-3dd2e51768a2" alt=""><figcaption></figcaption></figure>

A **product setup** might combine three: hub-and-spoke around a customer register with usage data and support tickets attached, a daisy chain following acquisition through activation to retention, and floating connections carrying shared themes such as onboarding, trust and pricing.

<figure><img src="https://2208645187-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FgczwnwqbMALFEXVtg0t1%2Fuploads%2F4287WLZyVbmtctdMqSRT%2Fimage.png?alt=media&amp;token=18d238f5-e24b-43bc-93a5-0b16cb5fca53" alt=""><figcaption></figcaption></figure>

A **research setup** often pairs a daisy chain with isolated datasets. A mixed-method study might begin with a demographic study, from which some of those users go on to take a survey, and a smaller group of them are then interviewed. Each stage draws its participants from the one before it, which is what makes it a chain. That same sequence may be running in several geographies at once though, with a different structure of questions in each, so the chains stay isolated from one another rather than rolling up into a single comparable set.

<figure><img src="https://2208645187-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FgczwnwqbMALFEXVtg0t1%2Fuploads%2FjH3iFPaeQqy7bWCKUDyZ%2Fimage.png?alt=media&amp;token=47c92df3-cdb4-4b45-a550-fca39ce5e005" alt=""><figcaption></figcaption></figure>

### Choosing the right data connector for your dataset

A few principles worth keeping in mind as you decide:

* **Start from the question, not the data.** Ask what you want to be able to say at the end of the project, then work backwards to the connector that makes it possible. One that looks tidy but cannot answer your question is not the right one.
* **Only connect what needs connecting.** Every link is something to maintain. If two datasets will never be analysed together, leaving them isolated is a legitimate choice rather than a missed opportunity.
* **Shared tags are a good place to begin.** They give you comparison across sources with far less setup than a register, and they do not stop you from adding record-level connections later.
* **Hub-and-spoke needs a real common entity.** If the thing at the centre is not consistently identified across your datasets, the register will not hold, and the connector will create more work than it saves.
* **Chain only where sequence matters.** A daisy chain is worth its dependencies when the order of events is part of your finding. Where it is not, a simpler connector will serve you better.
* **Connectors can change.** You are not locked in. Most workspaces start simple and add connections as the analysis matures and it becomes clearer which links are worth maintaining.
