Connector for Apache Spark

The Solace Connector for Apache Spark is a Spark DataSource V2 library that lets an Spark job read from and write to a Solace event broker. The Connector for Apache Spark supports both directions:

  • Source — Stream messages from a queue on a Solace event broker into a Spark DataFrame using spark.readStream.

  • Sink — Publish rows from a Spark DataFrame to a Solace event broker using writeStream or forEachBatch.

"Source" and "Sink" here describe direction relative to Spark, matching the Spark product's own terminology. This terminology is the reverse of the way Micro-Integrations use "source" and "target", which describe direction relative to the external system.

Unlike the Micro-Integrations described elsewhere in this section, the Connector for Apache Spark is not a standalone deployable service. Instead, you add it as a dependency to your own Spark application and use it directly from Spark code, for example:

spark.readStream.format("solace")
    .option("host", "tcp://<host>:55555")
    .option("vpn", "default")
    .option("username", "default")
    .option("password", "default")
    .option("queue", "my-queue")
    .load()

Supported Apache Spark Releases

This connector supports two release lines of Apache Spark. Configuration options and read/write behavior are identical between the two lines: only the runtime prerequisites and packaging differ.

  • The 4.x line is recommended for new deployments, and the 3.x line is necessary only if your environment is fixed at Apache Spark 3.5.x.

  • The 3.x line is supported until approximately September 2027.

  • For guidance on moving from the 3.x line to the 4.x line, see Migrating to the 4.x Line.

The following table compares the runtime requirements for each release line.

3.x Line 4.x Line

Apache Spark

3.5.x

4.0.x

Scala

2.12

2.13

Java

11

17

Maven coordinate

com.solacecoe.connectors:pubsubplus-connector-spark

com.solacecoe.connectors:pubsubplus-connector-spark

Before You Begin

Before you use the Connector for Apache Spark, ensure you meet these requirements:

  • You must have a Spark environment matching one of the release lines described in Supported Apache Spark Releases.

  • You must ensure the queue you intend to read from or write to already exists on the Solace event broker, with any needed topic subscriptions already added. The Connector for Apache Spark does not create this queue for you.

  • If you use the source (read) side, you must ensure the client username has consume permission on that queue.

  • If you use the source side, you must decide whether to let the connector auto-provision the checkpoint Last Value Queue (LVQ) or have an event broker administrator pre-create it. Either way, the client username's ACL must permit publish and subscribe access on the LVQ's topic. For more information, see Checkpointing.

  • If you plan to use message replay, you must ensure the Replay Log feature is enabled and configured on the event broker.

  • Only streaming reads are supported. This connector does not implement batch reads using spark.read.format("solace").

  • This connector does not support partitioned queues.

Configuring the Connector

To learn more about configuring and using the Connector for Apache Spark, see: