apache apex - hadoop users group

Pramod Immaneni <pramod@datatorrent.com>PPMC Member, Architect @DataTorrent IncDec 2nd, 2015

Stream Processing Architecture and ApplicationsApache Apex (incubating)

Apex Platform Overview

Apache Malhar Library

Native Hadoop Integration

• YARN is the resource manager

• HDFS used for storing any persistent state

Application Programming Model

A Stream is a sequence of data tuplesAn Operator takes one or more input streams, performs computations & emits one or more output streams

• Each Operator is YOUR custom business logic in java, or built-in operator from our open source library• Operator has many instances that run in parallel and each instance is single-threaded

Directed Acyclic Graph (DAG) is made up of operators and streams

Directed Acyclic Graph (DAG)

Filtered Stream

Output StreamTuple Tuple

Filtered Stream

Enriched Stream

Enriched

Stream

Operator

Advanced Windowing Support

Application window Sliding window and tumbling window

Checkpoint window No artificial latency

Application Specification

Partitioning and unification

Advanced Partitioning

Dynamic Partitioning

• Partitioning change while application is runningᵒ Change number of partitions at runtime based on statsᵒ Determine initial number of partitions dynamically

• Kafka operators scale according to number of kafka partitionsᵒ Supports re-distribution of state when number of partitions changeᵒ API for custom scaler or partitioner

unifiers not shown

1b 2c 3b

How tuples are partitioned• Tuple hashcode and mask used to determine destination

partitionᵒ Mask picks the last n bits of the hashcode of the tupleᵒ hashcode method can be overridden

• StreamCodec can be used to specify custom hashcode for tuplesᵒ Can also be used for specifying custom serialization

tuple: {Name, 24204842, San Jose}

Hashcode: 001010100010101

Mask (0x11)

Partition

Custom partitioning• Custom distribution of tuples

ᵒ E.g.. Broadcast

tuple:{Name, 24204842, San Jose}

Hashcode: 001010100010101

Mask (0x00)

Partition

Fault Tolerance• Operator state is checkpointed to a persistent store

ᵒ Automatically performed by engine, no additional work needed by operator

ᵒ In case of failure operators are restarted from checkpoint stateᵒ Frequency configurable per operatorᵒ Asynchronous and distributed by defaultᵒ Default store is HDFS

• Automatic detection and recovery of failed operatorsᵒ Heartbeat mechanism

• Buffering mechanism to ensure replay of data from recovered point so that there is no loss of data

• Application master state checkpointed

Processing GuaranteesAtleast once• On recovery data will be replayed from a previous checkpoint

ᵒ Messages will not be lostᵒ Default mechanism and is suitable for most applications

• Can be used in conjunction with following to ensure data is written once to store in case of fault recoveryᵒ Transactions with meta information, Rewinding output, Feedback from

external entity, Idempotent operationsAtmost once• On recovery the latest data is made available to operator

ᵒ Useful in use cases where some data loss is acceptable and latest data is sufficient

Exactly once• Operators checkpointed every window

ᵒ Can be combined with transactional mechanisms to ensure end-to-end exactly once behavior

Stream Locality• By default operators are deployed in containers (processes)

randomly on different nodes across the Hadoop cluster• Custom locality for streams

ᵒ Rack local: Data does not traverse network switchesᵒ Node local: Data is passed via loopback interface and frees up

network bandwidthᵒ Container local: Messages are passed via in memory queues

between operators and does not require serializationᵒ Thread local: Messages are passed between operators in a same

thread equivalent to calling a subsequent function on the message

Data Processing Pipeline ExampleApp Builder

Monitoring ConsoleLogical View

Monitoring ConsolePhysical View

Real-Time DashboardsReal Time Visualization

ResourcesApache Apex Community Page - http://apex.incubator.apache.org/

We Need Your Vote (Today)

Introducing Apache Apex - Not Just Another Stream Processing PlatformNext Gen Big Data Analytics with Apache ApexEnterprise-grade streaming under 2ms on Hadoop

Extra Slides

Partitioning and Scaling Out

• Operators can be dynamically scaled• Flexible Streams split• Parallel partitioning

• MxN partitioning • Unifiers

Fault Tolerance OverviewStateful Fault Tolerance Processing Semantics Data Locality

Supported out of the box– Application state– Application master state– No data loss

Automatic recovery Lunch test Buffer server

At least once At most once Exactly once

Stream locality for placement of operators

Rack local – Distributed deployment

Node local – Data does not traverse NIC

Container local – Data doesn’t need to be serialized

Thread local – Operators run in same thread

Data locality

Machine Data ApplicationLogical View

Machine Data ApplicationPhysical View

apache apex - hadoop users group

Software

apache hadoop 3 current status ajisaka -...

end-to-end "exactly-once" processing using apache apex (next...

nosql, apache solr and apache hadoop

machine learning support in apache apex (next gen hadoop)...

intro to apache apex (next gen hadoop) & comparison to spark...

introduction to apache apex

20100130 hadoop apache

abdw17-lightning talks track-security in a streaming...

developing streaming applications with apache apex (strata +...

real-time dashboards for apache apex (next gen hadoop) apps

apache hadoop

iot ingestion & analytics using apache apex - a native...

apache hadoop 1.1

apache hadoop india summit 2011 talk "making apache hadoop...

intro to apache apex, the next gen hadoop platform for...

integrating apache nifi and apache apex

apache hadoop ecosystem - lias (lab · apache hadoop...

introduction to apache hadoop & pig -...

apache apex (next gen hadoop) vs. storm - comparison and...

apache hadoop 2.0