Everything you need to know
Frequently Asked Questions
DataStori is an agentic platform that enables companies to get AI-ready. DataStori provides data and context for AI agents by extracting and unifying data from multiple cloud applications and enriching it with business goals and metrics.
This enables companies to rapidly and effectively deploy AI agents for data-driven decision making.
DataStori connects to a wide range of cloud applications like NetSuite, SAP Hana, ServiceTitan and HubSpot. It reads source documentation and creates data pipelines in real-time. This enables DataStori to connect to virtually any cloud application, unlike other tools that have a limited library of connectors.
DataStori is much more than a data connector or ingestion tool. It transforms and unifies the ingested data to create a common data layer with business context. In addition:
Functionality: DataStori creates data pipelines in real-time by reading source API documentation. Other tools have a library of connectors, with multi-step processes and lead times to add new ones. So, DataStori's customers can readily get their data from applications not served by other connectors.
Security: DataStori executes data pipelines in the customer's cloud. Data source and destination are both in the customer's cloud, and data never leaves their environment. This ensures that data handled by DataStori is always compliant with the customer's data security and privacy policies. Further, it eliminates the extra data hop and added cost that other connectors impose.
Pricing: DataStori's pricing is per connected application, not based on the number of pielines or ingested data volume. Because it runs serverless, DataStori spins up and shuts down infrastructure on-demand, making it transparent and cost-effective.
DataStori connects to the application API or database using the data owner's credentials, then creates and runs data pipelines from the source to the user-specified destination.
In addition to APIs and databases, DataStori can create pipelines from emailed csv files, SFTP folders and SharePoint.
A data pipeline is a component that copies data from a source to a destination via an integration. For example, a data pipeline can copy the NetSuite General Ledger table (source) to Azure SQL (destination). The integration specifies the copying and automation parameters including data deduplication, pipeline schedule, source columns, data backload and others.
DataStori orchestrates data pipelines from its cloud (AWS US East-1 region) but executes them in the customer's cloud. Data source and destination are both in the customer's cloud, and data never leaves their environment.
DataStori creates a Lakehouse in the customer's cloud and follows the Medallion architecture for data management. Files are written in the delta format and pushed to a data warehouse of the customer's choice, e.g., Azure SQL, Snowflake, PostgreSQL or any SQL Alchemy supported database.
Users consume the ingested data from the Lakehouse or the data warehouse. Besides delta, DataStori can store data in Iceberg, Parquet or csv formats in the Lakehouse.
Yes, DataStori performs a set of transformations on ingested data. It dedupes and flattens it and encrypts user-specified columns.
Further, DataStori unifies and enriches the data by enabling users to build semantic models and define business metrics across multiple data sources. This creates the common data layer (single pane of glass), which is the foundation for AI agents.
Users need to license and connect their own BI tools (e.g., Power BI) to the DataStori destination data store. Data visualization is not part of the DataStori subscription.
DataStori offers consulting services to help define business KPIs and design reports and dashboards.
Customers do not need to buy IT infrastructure to run DataStori. They need to have cloud licenses from AWS, MS Azure or GCP to provision servers and storage. These are directly licensed and paid for, and are not part of the DataStori subscription.
DataStori uses the customer's cloud licenses to spin up servers and other components to run pipelines. It shuts them down after pipeline execution, ensuring that provisioning is on-demand, with minimal fixed costs when no pipelines are running.
DataStori scales server and storage infrastructure on-demand. It has run data pipelines on tables as large as 100 GB, with 30 million rows. Pipelines can be configured and scheduled to backload multi-year data.
The only constraints to data size are the rate limits in the source APIs and database connections. Breaking a data load into smaller datasets addresses this issue.
Users can set up and run as many data pipelines as they want. The number of data pipelines is only constrained by the user's cloud services provider.
DataStori charges customers by the number of application instances connected. This fee has two components: one-time setup and monthly licensing.
DataStori does not charge users on the volume of data ingested or the number of pipelines created and executed. These are part of the infrastructure cost that customers directly pay their cloud services provider.
DataStori is SOC 2 Type 2 compliant.
DataStori orchestrates data pipelines from its cloud (AWS US East-1 region) but executes them in the customer's cloud. Data sources and destinations are in the customer's cloud, ensuring compliance with the customer's data policies at all times.
Further, DataStori can encrypt user-specified columns from a data source or drop them from the final output.
Yes, users need to allow DataStori to access their cloud infrastructure to spin up servers and other components. DataStori also needs access to the source application APIs or databases from where data is to be ingested.
All user credentials are secure in DataStori. Application API tokens are encrypted using AES 256 and stored in the application database. They cannot be read by the DataStori admin or anyone else.
DataStori doesn't need any credentials to the destination data store because the required permissions are assigned to the servers spun up for pipeline execution.
Security elements in DataStori include data encryption, virtual network, multi-factor authentication, alerts and logging.
DataStori cannot view business data. While DataStori orchestrates data pipelines from its cloud, the data movement from source application to storage destination is entirely in the customer's cloud. DataStori can only create and access the metadata for pipeline setup and execution.
DataStori's AI agent identifies relevant datasets, auto-documents them and joins data from multiple sources to create a unified view. This common data layer serves as the foundation for AI agents.
DataStori runs the following checks on all ingested data for every pipeline execution:
- Data freshness test, to check when the data was last refreshed
- Primary key not null test
- Primary key uniqueness test, to ensure that there are no duplicates in the primary key.
In addition, DataStori has automated retries, logging and alerts to make data pipelines reliable and robust.
By default, data pipelines in DataStori have a concurrency of 1. Only one instance of a pipeline can run at a given time, and all other triggered instances are queued. In addition, output data is saved in the delta format and supports ACID compliance.
DataStori auto-documents tables and columns into a data catalog, enabling automated schema evolution and tracking. In addition, data and schema changes can be rolled back to a defined restore point if required.
DataStori supports the following API authentication protocols:
1. API Key
2. Basic Authentication
3. OAuth2: Client Credentials and Authorization Grant
In addition, DataStori can be extended to support custom authentication flows.
Have more questions?
See our product documentation or get in touch with our team