Thank you for Subscribing to CIO Applications Weekly Brief
Thank you for Subscribing to CIO Applications Weekly Brief
IData has been recognized by CIO Applications Magazine as the exclusive recipient of “Top 10 Amazon Cloud Solutions Companies - 2022,” based on our proprietary methodology, reflecting its position in the industry, and is also named among “,” reflecting its broader leadership. This profile has been developed by the CIO Applications research and editorial team based on insights from an interview with Todd Fearn, Founder and Managing Partner.
Todd Fearn, Founder and Managing PartnerIn a conversation with CIO Applications magazine, Todd Fearn, Founder and Managing Partner of IData, explains how the company offers a centralized solution that runs on JSON configurations and not handwritten code.
Could you shed light on your cloud-based offering?
The IData pipeline is designed to offer standardized dataset registration, ingestion, validation, format conversion, and consumption patterns according to client’s requirements and can be spun up in an hour or less using our Infrastructure as Code (IaC), written entirely in Terraform. It runs on the client’s AWS account and allows them to feed their datasets into it and the datasets are automatically validated and transformed into a data lake or a data warehouse like Snowflake or Redshift. Our solution runs on JSON-based dataset configurations and is designed using a microservices architecture, leveraging Delta Lake open-source for Lakehouse capabilities—including upsert support with fast, pluggable indexing, data publishing with rollback support, synch compaction, and time travel querying.
Our pipeline can help eliminate over 12 months’ worth of PoC and eventual data technology implementation. Most importantly, clients need not configure and maintain the servers as the solution is built, using IaC, and runs in a serverless environment. Currently, the pipeline has thousands of datasets flowing through it daily at major financial institutions.
What methodology do you use to help clients get started with the IData pipeline?
Most of our clients find it difficult to build an efficient and easy-to-monitor pipeline by leveraging some of the tools in the cloud.
On the contrary, we help clients move to our centralized cloud-based platform where they can easily register their datasets using our API. Following this, our pipeline ingests all the data, allowing us to validate and cleanse it and deliver to the data warehouse or data lake, based on the configuration of the client’s choice. The validation is done using data quality configuration rules and data deduplication capabilities. Following this, the datasets are automatically converted to an Apache Parquet format for low-latency consumption in S3.
This data is then available for consumption in the downstream systems.
Our pipeline is also equipped with a notification system using an SQS or SNS message that alerts downstream consumers when the data is available to retrieve and process. As all data status flows into our dataset monitoring environment, clients can easily track their datasets through the pipeline in real-time, using our user interface. They can view the name of the dataset and access information regarding when the data landed, when the process was completed, how long it took, and if it succeeded or failed to reach a data warehouse or data lake. With this visibility, clients can assess the reason for a particular data landing failure and automatically trigger a support team via email for proactive resolution.
Could you elaborate on a particular client success story?
A large foreign country pension fund had several investment groups that wished to move from fundamental to analytical investment strategies by leveraging big data and their internal system data. However, the existing internal system data was scattered across many on-premise databases, and the big data sets were yet to be acquired from third parties. This gave rise to a need for a robust solution to ingest their internal and external data into a centralized data store.
How is IData a cut above the rest in the industry?
We believe in delivering an end-to-end data pipeline where clients can plug in our solution and get up and running seamlessly in a few days. Another aspect where we outshine the others is that we can assist clients who do not want to write their own configurations. In such cases, we provide our UI where they can leverage our wizard that automatically generates the JSON configuration. Our goal ultimately is to offer a convenient, seamless, and independently deployable pipeline that organizations can leverage to securely access a single source of truth.
CIO Applications Weekly Brief
Be first to read the latest tech news, Industry Leader's Insights, and CIO interviews of medium and large enterprises exclusively from CIO Applications
