Categories: AnalyticsBusiness IntelligenceHadoopIn Memory

Comparison of Stream Processing Frameworks and Products

See how products, libraries, and frameworks that full under ‘streaming data analytics’ use cases are categorized and compared.

Streaming Analytics processes data in real time while it is in motion. This concept and technology emerged several years ago in financial trading, but it is growing increasingly important these days due to digitalization and Internet of Things (IoT). The following slide deck from a recent talk at a conference covers:

Real world success stories from different industries (Manufacturing, Retailing, Sports)
Alternative Frameworks and Products for Stream Processing
Complementary Relationship to Data Warehouse, Apache Hadoop, Statistics, Machine Learning, Open Source R, SAS, Matlab, etc.

Stream Processing Frameworks and Products

The following picture shows the key differences between frameworks (no matter if open source such as Apache Storm, Apache Flink, Apache Spark or closed source such as Amazon Kinesis) and products (such as TIBCO StreamBase / Live Datamart, IBM InfoSphere Streams, Software AG’s Apama).

Of course, you can implement everything by writing code and using one or more frameworks. However, besides several other benefits, the key differentiator of using a product is time to market. You can realize projects in weeks instead of months or even years. Delivering quickly is the number one priority of most enterprises these days in a world where the only constant is change!

I recommend that you choose one or two frameworks and one or two products to implement a proof of concept (POC); spend e.g. five days with each one to implement a streaming analytics use case, which includes integration of input feeds or sensors, correlation / sliding windows / patterns, simulation and testing, and a live user interface to monitor and act proactively. At the end, you can compare the results and decide which fits you best.

Fast Data and Streaming Analytics in the Era of Hadoop, R and Apache Spark

The following slide deck discusses the above topics in much more detail:

Click on the button to load the content from www.slideshare.net.

Load content

Parts of this (extensive) slide deck were used for talks at several international conferences such as JavaOne 2015 in San Francisco. I appreciate any feedback about the content to improve it continuously…If you want to learn more about Streaming Analytics and its relation to Big Data and Apache Hadoop, I recommend the following InfoQ article: Real-Time Stream Processing as Game Changer in a Big Data World with Hadoop and Data Warehouse.

Kai Waehner

bridging the gap between technical innovation and business value for real-time data streaming, processing and analytics

Next Microservices = Death of the ESB? (2016, Meetup Dublin) »

Previous « Difference between a Data Warehouse and a Live Datamart?

Published by

Kai Waehner

10 years ago

How Penske Logistics Transforms Fleet Intelligence with Data Streaming and AI

Real-time visibility has become essential in logistics. As supply chains grow more complex, providers must…

1 day ago

SAP Sapphire

Data Streaming Meets the SAP Ecosystem and Databricks – Insights from SAP Sapphire Madrid

SAP Sapphire 2025 in Madrid brought together global SAP users, partners, and technology leaders to…

6 days ago

Agentic AI

Agentic AI with the Agent2Agent Protocol (A2A) and MCP using Apache Kafka as Event Broker

Agentic AI is emerging as a powerful pattern for building autonomous, intelligent, and collaborative systems.…

1 week ago

Fantasy Sports

Powering Fantasy Sports at Scale: How Dream11 Uses Apache Kafka for Real-Time Gaming

Fantasy sports has evolved into a data-driven, real-time digital industry with high stakes and massive…

2 weeks ago

Lakehouse

Databricks and Confluent Leading Data and AI Architectures – What About Snowflake, BigQuery, and Friends?

Confluent, Databricks, and Snowflake are trusted by thousands of enterprises to power critical workloads—each with…

3 weeks ago

Databricks and Confluent in the World of Enterprise Software (with SAP as Example)

Enterprise data lives in complex ecosystems—SAP, Oracle, Salesforce, ServiceNow, IBM Mainframes, and more. This article…