[ Tech stack ]
Apache Spark
The distributed analytics engine for processing terabytes.
Spark processes massive data volumes in memory across a distributed cluster. Its APIs (DataFrame, SQL, Streaming, MLlib) cover batch, stream and machine learning. Paired with Delta Lake, it's the foundation of modern lakehouses.
[ Why Apache Spark at Dexon ]
What this technology does well, and why we use it.
Typical usage: Large-volume ETL, lakehouse, distributed ML training.
- 01
Distributed in-memory processing: orders of magnitude vs Hadoop.
- 02
Python, Scala, SQL, R APIs: mixed teams accommodated.
- 03
Structured Streaming for near-real-time.
- 04
Databricks, EMR, Dataproc: managed on all three clouds.
[ Complementary technologies ]
The building blocks we often mobilise alongside.
A stack rarely exists alone. Here are the technologies Dexon most often pairs with this one, through pipeline habits, usage similarity or internal mastery. Click on a building block to see its scope.
[ Reassurance ]
- 0+
- custom projects delivered
- 30
- engineers, designers, project managers
- 80 %
- from top French schools
- 24 h
- average reply time
[ They trust us ]
More than 100 French and European companies trust us
















[ Press ]
They talk about us.
Application and data division.
Nationwide coverage by BFM Business, Le Figaro, Challenges, La Tribune and CNews. An outside reading of our work and our innovations.


