StreamSets Data Collector is an open source data integration tool that can ingest data from various sources in both batch and streaming modes. It uses a record-oriented approach to data processing which avoids issues caused by combinatorial explosion. Pipelines can be developed visually using an IDE interface, allowing non-technical users to build integrations. StreamSets originated from ex-Cloudera and Informatica employees and focuses on continuous open source development.