> For the complete documentation index, see [llms.txt](https://dots.gitbook.io/dots-docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://dots.gitbook.io/dots-docs/more-about-dots/data-security-and-privacy-policy.md).

# Data Security & Privacy Policy

### **1. Introduction**&#x20;

Dots is a qualitative data management platform. It enables teams to collect, organize, and analyze qualitative data through a single, unified platform. Dots serves organizations across sectors including international development, public health, education, and social impact research.&#x20;

This Data Security and Privacy Policy outlines the measures, principles, and practices that Dots employs to safeguard the confidentiality, integrity, and availability of all data processed through its platform. Given the sensitive nature of qualitative research data—which may include personal narratives, field observations, survey responses, interview transcripts, and community-level insights—Dots is committed to maintaining the highest standards of data protection.&#x20;

This document is intended for Dots customers, prospective clients, partners, auditors, and any stakeholders who require transparency into how Dots handles data security and privacy.&#x20;

#### What is Dots?&#x20;

Dots is a multi-tenant SaaS platform for qualitative and quantitative data management, analysis, and reporting. The platform enables organizations to:&#x20;

* Collect data through forms, surveys, and structured data entry&#x20;
* Import data from external sources (CSV/Excel files, Slack, Fathom, Sarvam audio transcription)
* Organize and annotate data with thematic tagging and taxonomies&#x20;
* Analyze data using AI-powered tools (chat-based Q\&A, automated annotation, summarization, report generation)&#x20;
* Visualize and export findings&#x20;

Each client organization operates within an isolated tenant environment with its own database, configurations, and access controls.&#x20;

### **2. Scope**&#x20;

This policy applies to all data collected, stored, processed, transmitted, and managed through the Dots platform, including:&#x20;

**Qualitative research data:** interview transcripts, field notes, survey responses, observational records, audio and video transcripts, photographs, and any other unstructured or semi-structured data inputs.&#x20;

**User account data:** information associated with platform user accounts, including names, email addresses, organization affiliations, roles, and access credentials.&#x20;

**Metadata:** system-generated data such as timestamps, geolocation tags, device information, session logs, and audit trails.

**Analytical outputs:** AI-generated summaries, coded themes, annotations, tagged datasets, and insights reports produced within the platform.&#x20;

This policy covers all environments in which Dots operates, including production, staging, and development environments, as well as any third-party integrations connected to the platform.&#x20;

### **3. Cloud Infrastructure & Deployment**&#x20;

#### 3.1 Hosting Provider&#x20;

The entire OKF platform is deployed on Google Cloud Platform (GCP), under project ID ok framework . Google Cloud maintains comprehensive security certifications including:&#x20;

* SOC 2 Type II — regular third-party audits of security controls&#x20;
* ISO/IEC 27001:2022 — information security management&#x20;
* ISO/IEC 27017 — cloud security controls&#x20;
* ISO/IEC 27018 — protection of personal data in the cloud&#x20;
* GDPR compliance — data processing agreements and appropriate controls&#x20;

For details, see *Google Cloud Compliance Offerings*.&#x20;

#### 3.2 Compute & Application Hosting&#x20;

Our application services run on Google App Engine Standard Environment. All App Engine services are configured with secure: always , meaning all traffic is served exclusively over HTTPS. Google App Engine enforces TLS 1.2+ for all connections.&#x20;

#### 3.3 Database&#x20;

* **MongoDB** — Our primary database. Connections are managed through a centralized connector with per-tenant isolation. Connection credentials are stored in Google Cloud Secret Manager (never in code).&#x20;
* **Elasticsearch** — Used for full-text search and vector embeddings. Cloud-hosted with API key authentication, credentials stored in Secret Manager.&#x20;
* **Redis** — Used for in-memory caching of templates and platform configurations. Credentials stored in Secret Manager.&#x20;

#### 3.4 File Storage&#x20;

Uploaded media and files are stored in Google Cloud Storage (GCS).&#x20;

* Buckets are organized per-tenant (either dedicated per-tenant buckets or tenant-prefixed folders within shared environment buckets).&#x20;
* Files are encrypted at rest using Google-managed encryption keys (AES-256) by default
* GCS service account credentials are stored in Secret Manager&#x20;

#### 3.5 Secret Management&#x20;

All sensitive credentials (database URIs, API keys, JWT secrets, service account keys) are stored in **Google Cloud Secret Manager** and retrieved at application startup. Secrets are never committed to source code or stored in plaintext configuration files.&#x20;

#### 3.6 Inter-Service Communication&#x20;

* Google Cloud Pub/Sub is used for real-time cache synchronization across server instances and for event-driven processing between services&#x20;
* Internal service authentication uses dedicated JWT tokens for backend-to-backend communication between okf-be and okf-sub&#x20;

#### 3.7 Error Monitoring & Observability&#x20;

* Sentry is used for distributed error tracking and performance monitoring&#x20;
* Google Cloud Logging is used for infrastructure-level logs&#x20;
* OpenTelemetry is used in okf-sub for distributed tracing&#x20;

### **4. Categories of Data**&#x20;

We process two fundamentally distinct categories of data:&#x20;

#### Category A — Client Data (Your Data)&#x20;

This is data that **you bring into the platform** or that **your users create**. This includes:&#x20;

* User accounts and profiles (names, emails, phone numbers)&#x20;
* Content created through forms and surveys&#x20;
* Data imported from files (CSV, Excel) or external integrations (Slack, Fathom, etc.)
* Uploaded media files (images, documents, audio, video)&#x20;
* Tags, annotations, and thematic coding&#x20;
* Reports and analytical outputs&#x20;

**This data belongs to you.** We process it solely to provide the services you have contracted for. We do not sell, share, or use your data for any purpose other than delivering and operating the platform on your behalf.&#x20;

#### Category B — Platform-Collected Data (Our Data)&#x20;

This is data that **we collect** for the purposes of product improvement, analytics, error monitoring, and internal operations monitoring. This includes:&#x20;

* Product analytics events (via Mixpanel)&#x20;
* Web analytics (via Google Analytics)&#x20;
* Session recordings and heatmaps (via Hotjar, where enabled)
* Error and performance data (via Sentry)&#x20;
* Aggregated tenant usage statistics.&#x20;

### **5. Platform-Collected Data — Data We Collect**&#x20;

This section describes the data that **Dots collects** about your usage of the platform. We want to be fully transparent about what is collected, by which tools, and for what purposes.&#x20;

#### 5.1 Mixpanel — Product Analytics&#x20;

**What it is**: Mixpanel is a product analytics service that helps us understand how users interact with the platform so we can improve the product.&#x20;

**What data is sent to Mixpanel:**&#x20;

User identification:&#x20;

* User ID (internal database ID)&#x20;
* User name&#x20;
* User email address&#x20;
* User role&#x20;

Tenant properties (sent with every event):&#x20;

* Tenant ID&#x20;
* Environment (dev/staging/prod)&#x20;

Events tracked:&#x20;

* Login flow: Page viewed, form engaged, login attempted (with method: email/phone), OTP sent, login success/failure (with failure reason). For phone login, a masked phone number is sent (not the full number).&#x20;
* Language selection: Language chosen, time spent&#x20;
* Discovery filters: Filter selected/deselected with filter details&#x20;
* Search & sort: When search or sort is used&#x20;
* Annotation explorer: Content template selections, analysis type tab switches
* Card interactions: Card clicks, tag clicks, card expansions&#x20;

**What is NOT tracked:** The actual content of your data (survey responses, imported documents, etc.) is never sent to Mixpanel.&#x20;

**Mixpanel's data practices:** Mixpanel acts as a data processor. Data is stored in Mixpanel's infrastructure (with options for EU or India data residency). Mixpanel retains event data for 2–5 years depending on account creation date. See *Mixpanel Privacy Policy and Mixpanel GDPR Compliance.*&#x20;

#### 5.2 Google Analytics (GA4) — Web Analytics&#x20;

What it is: Google Analytics 4 is used for web traffic analytics and pageview tracking. Configuration:&#x20;

* GA is only active in the production environment
* GA can be enabled or disabled per tenant via platform configuration (default: disabled)&#x20;

**What data is collected:**&#x20;

* Pageview data (URL paths)&#x20;
* Standard web analytics (browser, device, geography, referrer — as collected by GA4 by default)
* Internal traffic is flagged: users with @ooloilabs.in email addresses are tagged as traffic\_type: "internal"&#x20;

**What is NOT collected:** No custom events beyond pageviews are currently tracked via GA. No content data is sent.&#x20;

**Google's data practices:** GA4 data retention is configurable (2–26 months for user-level data, up to 14 months by default for standard users). Google provides data deletion mechanisms for individual users. See GA4 Data Retention and Google Privacy Policy.&#x20;

#### 5.3 Hotjar — Session Recordings & Heatmaps&#x20;

**What it is:** Hotjar captures session recordings (replays of user interactions) and heatmaps to help us understand how users navigate the platform.&#x20;

**What data is collected:**&#x20;

* Mouse movements, clicks, scrolling behavior&#x20;
* Page navigation sequences&#x20;
* Device type, screen size, browser, geographic location, preferred language&#x20;
* IP address (anonymized by Hotjar by default)&#x20;

**What is NOT collected:** Hotjar suppresses sensitive input fields by default (credit card numbers, passwords). Hotjar does not collect or sell personal data.&#x20;

**Hotjar's data practices:** Hotjar acts as a data processor. Data is hosted on AWS. Hotjar does not sell personal data. Hotjar is GDPR-compliant and provides a Data Processing Agreement. See Hotjar Privacy Policy and Hotjar GDPR Commitment.&#x20;

#### 5.4 Sentry — Error & Performance Monitoring&#x20;

**What it is:** Sentry captures application errors, exceptions, and performance data to help us diagnose and fix bugs.&#x20;

**What data is sent to Sentry:**&#x20;

* Error stack traces and exception details&#x20;
* User email address&#x20;
* User name&#x20;
* IP address (auto-captured by Sentry)&#x20;
* Browser, device type, and OS information&#x20;

**What is NOT collected:** Sentry does not receive your content data, survey responses, or imported documents. It receives only technical error information and the minimal user context needed to diagnose issues.

**Sentry's data practices:** Sentry collects only the data you configure to be sent. Sentry does not use your data to track users. *See Sentry Privacy Policy.*&#x20;

#### 5.5 Activity Analytics — Internal Usage Logging&#x20;

**What it is:** An internal system that logs user activity within the platform, stored in your tenant's own MongoDB database (the activityAnalytics collection).&#x20;

**What is logged:**&#x20;

* Activity type (content created, updated, deleted, published, moderated, AI annotation applied, etc.)&#x20;
* User who performed the action (ID, name, avatar)&#x20;
* Content that was acted upon (content type, ID, title)&#x20;
* Timestamp&#x20;

**Purpose:** This data is used to power activity feeds and audit trails within the platform. It remains within your tenant's database and is subject to the same access controls as your other data.&#x20;

#### 5.6 Tenant Usage Statistics&#x20;

**What it is:** An internal administrative system that aggregates high-level usage metrics across tenants and manages them to an internal Dots tenant for operational monitoring. This is triggered manually (not automatic) and requires internal service authentication.&#x20;

**What is aggregated:**&#x20;

| **Metric**                                                                               | **Description**                                                                                                |
| ---------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- |
| Number of data sets                                                                      | Count of content types per tenant.                                                                             |
| <p>Total documents </p><p>Total fields </p><p>Estimated character/word/token counts </p> | <p>Count of documents across data sets Count of fields in all documents. </p><p>Volume Metrics for Content</p> |
| Taxonomy Metrics                                                                         | Count of themes, sub-themes, tags, and tag applications                                                        |
| AI Usage Metrics                                                                         | Counts of AI suggestions, annotations applied.                                                                 |

Number of reports/widgets Count of Created Reports Platform Configuration Snapshots Roles, deployment, and AI configuration&#x20;

**What is NOT aggregated:** Individual documents, user personal data, actual content text, or file attachments are not copied. Only aggregated counts and configuration metadata are collected.

The platform context fields (which are text descriptions of your organization's research goals that you configure in the platform) are aggregated as part of configuration metadata. These may contain descriptions of your organization's work.&#x20;

**Who has access:** Only Dots internal staff with @ooloilabs.in email accounts can access the mothership tenant. Access is enforced at the middleware level — all API requests to the analytics tenant (except login) require a JWT token from a user with an @ooloilabs.in email domain.&#x20;

### **6. AI Systems & Third-Party Model Providers**&#x20;

Dots integrates AI capabilities powered by third-party large language model (LLM) providers. This section describes what data flows to which providers and their respective data policies.&#x20;

#### 6.1 OpenAI — Primary AI Provider&#x20;

**What data is sent to OpenAI:**&#x20;

For chat-based Q\&A and analysis:&#x20;

* User's question/query&#x20;
* Conversation history (previous Q\&A pairs in the session)&#x20;
* Data schema (field names, types, labels, and sample values from your content types)&#x20;
* Platform context (your configured research goals/domain description)
* Tool call results (aggregated metrics, document excerpts retrieved from your database)&#x20;

For document embeddings (vector representations):&#x20;

* Text content from your documents (chunked text fields)&#x20;
* Annotation text excerpts&#x20;
* Document summaries&#x20;

For AI annotation/categorization:&#x20;

* Document text content (the specific text being annotated)
* Tag taxonomy (your themes, sub-themes, and tags)
* Platform context&#x20;

For document summarization:&#x20;

* Document text (truncated to token limits)
* Platform context&#x20;

**OpenAI's data policy (for API usage):**&#x20;

* **OpenAI does NOT use API data to train models** (since March 1, 2023). Data sent through the API is not used to train or improve OpenAI's models unless you explicitly opt in (we have not opted in).&#x20;
* **Data retention:** API data is retained for up to **30 days** for abuse monitoring purposes, then deleted.
* Eligible customers can request Zero Data Retention (ZDR).&#x20;

See *OpenAI Data Controls and OpenAI Data Usage Policy.*&#x20;

#### 6.2 Anthropic (Claude) — Secondary AI Provider \[BETA Features]&#x20;

**What data is sent to Anthropic:**&#x20;

For report generation:&#x20;

* User's report request/query&#x20;
* Data schema and platform context
* Report plan (structured plan generated in a prior step) Gathered data (aggregated metrics and retrieved data)&#x20;

For pattern analysis:&#x20;

* Statistical analysis of annotations (tag prevalence, co-occurrence, lift scores)&#x20;
* Aggregated statistics only — not individual document data
* Platform context for grounding insights&#x20;

**Anthropic's data policy (for API usage):**&#x20;

A**nthropic does NOT use API data to train models** under commercial terms, unless the customer explicitly opts in (we have not opted in).&#x20;

**Data retention:** API logs are retained for **7 days** (as of September 15, 2025), then fully deleted within 30 days.&#x20;

See *Anthropic Privacy Center.*&#x20;

#### 6.3 Sarvam AI — Speech-to-Text Provider \[Add-on Only]&#x20;

**What data is sent to Sarvam:** Audio files uploaded by your users (downloaded from GCS and sent as multipart form data) Audio is processed for speech-to-text transcription and optionally English translation with speaker diarization&#x20;

**Sarvam's data policy:** Sarvam collects the text, queries, and inputs you send and the responses returned.&#x20;

**Sarvam reserves the right to use input and output data to train, improve, and further develop their product.** Users explicitly consent to this by using the service. This is a notable distinction from OpenAI and Anthropic, which do not use API data for training by default.&#x20;

See *Sarvam AI Privacy Policy and Sarvam AI Terms of Use.*&#x20;

**Important note for clients:** If you use the Sarvam audio transcription import feature, the audio data you upload will be sent to Sarvam's API, and per their current terms, Sarvam may use that data to improve their models. If this is a concern, please discuss alternative transcription options with us.

#### 6.4 ElevenLabs — Speech-to-Text Provider \[Add-on, Emerging]&#x20;

**What data is sent to ElevenLabs:** Audio files uploaded by your users (downloaded from GCS and sent via ElevenLabs' API). Audio is processed for speech-to-text transcription and, where enabled, translation. ElevenLabs is being introduced as an alternate option to Sarvam for these features.&#x20;

**ElevenLabs' data policy:** By default, ElevenLabs' general privacy policy permits the company to use voice, text, and related data to research, develop, and improve its AI models.&#x20;

**Dots has opted out of this model-training use.** Audio and text sent to ElevenLabs as part of Dots' services is not used by ElevenLabs to train or improve its models. This is a notable distinction from Sarvam, which by default reserves the right to use input/output data for model improvement.&#x20;

See *ElevenLabs Privacy Policy.*&#x20;

**Important note for clients:** Dots is finalizing its commercial terms with ElevenLabs and will update this policy to confirm any additional data-handling details as this integration goes live.&#x20;

#### 6.5 How AI Data is Stored Locally&#x20;

After AI processing, results are stored in your tenant's database.&#x20;

| **Collection**                                   | **What is stored?**                                                                                                                                                                                             |
| ------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| <p>AI Chat </p><p>SYSTEM\_AISuggestionCache </p> | <p>Chat conversation history (user queries and AI responses), related document references, query metadata </p><p>Cached AI annotation suggestions per document field, with tag mappings and justifications.</p> |
| embeddings (MongoDB)                             | Text chunks with 1,536-dimensional vector embeddings.                                                                                                                                                           |

*chunks / annotations (Elasticsearch) Vector embeddings indexed for semantic search*&#x20;

*SYSTEM\_DocumentSummary AI-generated document summaries with hash-based invalidation*&#x20;

AI chat conversations are scoped to the user who created them — other users cannot access your chat history (enforced via user ID matching on read queries).&#x20;

#### 6.6 AI Data Access Control&#x20;

* Only content types that are explicitly enabled in your AI configuration are included in AI processing
* Field-level filtering ( includeFields / excludeFields ) allows you to control which fields are visible to AI systems&#x20;
* ACL-based scoping restricts what data appears in AI responses based on user permissions&#x20;

### **7. Access Control and User Management**&#x20;

#### 7.1 Role-Based Access Control (RBAC)&#x20;

Dots implements a granular role-based access control system that allows organizations to define who can view, edit, annotate, export, or administer data. Roles can be customized to match organizational hierarchies and project-specific requirements. Access permissions are enforced at the platform level, ensuring that users can only interact with data they are authorized to access.&#x20;

#### 7.2 Authentication&#x20;

User authentication is managed through secure credential systems. Dots supports password-based authentication with enforced complexity requirements. The platform supports integration with enterprise identity providers for single sign-on (SSO) capabilities where applicable.&#x20;

#### 7.3 Session Management&#x20;

Active user sessions are subject to configurable timeout policies. Sessions are invalidated upon logout, and inactive sessions expire automatically. Session tokens are generated using cryptographically secure methods and are resistant to prediction and replay attacks.&#x20;

#### 7.4 Audit Logging&#x20;

All user activities within the Dots platform are logged. Audit logs capture information including the user identity, action performed, timestamp, affected data record, and originating IP address. These logs are retained for a defined period and are available to authorized administrators for compliance and forensic review.&#x20;

### **8. Third-Party Integrations and Data Sharing**&#x20;

#### 8.1 Integration Security&#x20;

Dots supports integrations with external applications and services. All third-party integrations are vetted for security compliance before they are made available on the platform. Data exchanged through integrations is transmitted over encrypted channels, and API keys and authentication tokens are stored securely.&#x20;

#### 8.2 Data Sharing&#x20;

Dots does not share customer data with third parties except where explicitly authorized by the customer or required by law. When data sharing is necessary for service delivery (for example, with a cloud hosting provider), Dots ensures that appropriate data processing agreements are in place with the third party.

#### 8.3 Sub-Processors&#x20;

A list of sub-processors used by Dots is available upon request. Customers are notified of material changes to the sub-processor list, and Dots ensures that all sub-processors are contractually bound to equivalent data protection standards.&#x20;

### **9. Data Retention and Deletion**&#x20;

#### 9.1 Retention Policy&#x20;

Dots retains customer data for as long as the customer’s account remains active and as required for the purposes for which the data was collected. Organizations can configure retention periods for specific data categories within the platform.&#x20;

#### 9.2 Data Deletion&#x20;

Upon termination of a customer’s subscription or upon a valid deletion request, Dots will permanently delete or anonymize all customer data within a defined timeframe (typically 30 days), unless a longer retention period is required by law. Deletion is performed across all systems, including primary databases, backups, and logs.&#x20;

#### 9.3 Data Export&#x20;

Customers can export their data from the Dots platform at any time in standard, machine-readable formats. Dots supports full data portability to ensure that organizations are not locked into the platform.&#x20;

### **10. Incident Response**&#x20;

#### 10.1 Incident Detection&#x20;

Dots employs continuous monitoring and alerting systems to detect potential security incidents. Automated threat detection tools monitor for unauthorized access attempts, anomalous data patterns, and system vulnerabilities.&#x20;

#### 10.2 Response Procedure&#x20;

In the event of a confirmed security incident, Dots follows a structured incident response procedure that includes immediate containment, root cause analysis, remediation, and post-incident review. Affected customers are notified promptly in accordance with contractual obligations and applicable legal requirements.&#x20;

#### 10.3 Notification Timeline&#x20;

Dots commits to notifying affected customers of a confirmed data breach within 72 hours of discovery, consistent with GDPR requirements and industry best practices. Notifications include a description of the incident, the data affected, measures taken, and recommended actions for the customer.

### **11. Internal Security Practices**&#x20;

#### 11.1 Personnel Security&#x20;

All Dots team members with access to customer data or production systems are bound by confidentiality agreements. Access to production environments is limited to authorized personnel on a need-to-know basis.&#x20;

#### 11.2 Security Training&#x20;

Dots conducts regular security awareness training for all employees. Training covers topics including data handling procedures, phishing awareness, incident reporting, and secure development practices.&#x20;

#### 11.3 Secure Development&#x20;

The Dots engineering team follows secure software development lifecycle (SDLC) practices. Code changes undergo peer review, and security testing is integrated into the continuous integration and deployment pipeline. Vulnerability assessments and penetration testing are conducted periodically.&#x20;

### **12. Policy Updates**&#x20;

This policy is reviewed and updated at least annually, or more frequently as required by changes in law, technology, or business operations. Material changes to this policy will be communicated to customers through the Dots platform and via email notification. Continued use of the platform following notification of changes constitutes acceptance of the revised policy.&#x20;

### **13. Contact Information**&#x20;

For questions, concerns, or requests related to data security and privacy, please contact the Dots team:&#x20;

***Website:** <https://getdots.in>*&#x20;

***Contact:** <hello@getdots.in>*

<br>
