Troubleshooting
The following troubleshooting tips might help you resolve issues with the Solace Connector for Apache Spark. If problems persist, contact Solace.
- Batch Read Returns No Data or Fails
-
Only streaming reads are supported. If you call
spark.read.format("solace")instead ofspark.readStream.format("solace"), the read fails. UsereadStreamfor the source side. - SolaceInvalidPropertyException at Startup
-
-
The exception message
"Please provide Solace Queue in configuration options"means thequeueoption is missing. Add it before starting the source. -
The exception message
"Please set batch size greater than zero"means thebatchSizeparameter was set to a negative value. The connector rejects a negative value rather than treating it as unlimited; use0for no per-partition limit or a positive integer to cap how many messages each partition reads per micro-batch.
-
- NoSuchMethodError, ClassNotFoundException, or Similar Errors at Startup
-
These typically indicate a version mismatch between the connector, Spark, and Scala, or a stale JAR file left on the classpath:
-
Confirm that the connector version matches your Spark/Scala/Java baseline. See Supported Apache Spark Releases.
-
Remove any earlier connector JAR file from the classpath before installing a new version, and restart your Spark environment. This is especially important if you move across the 3.x/4.x boundary, because a leftover 3.x (Scala 2.12) JAR file is binary-incompatible with the 4.x line's Scala 2.13 runtime and causes exactly this kind of error.
-
On Databricks, a Maven library install failing outright (rather than the job failing at runtime) on a Standard or Shared access-mode cluster usually means the coordinate isn't on the workspace's Unity Catalog allowlist. Ask a metastore admin to add it, or use a Dedicated access-mode cluster instead.
-
- Connection Errors
-
If the job fails to connect to the event broker:
-
Verify that
host,vpn,username, andpassword(or the equivalent client certificate or OAuth options) are correct. -
Confirm that the client username's ACL permits connecting to the specified Message VPN.
-
Increase
connectRetriesandreconnectRetriesif the event broker is temporarily unavailable during job startup.
-
- OAuth Authentication Failures
-
If OAuth authentication fails:
-
Confirm that
solace.apiProperties.AUTHENTICATION_SCHEMEis set toAUTHENTICATION_SCHEME_OAUTH2. -
For server-fetch mode, verify that
solace.oauth.client.auth-server-url,client-id, andcredentials.client-secretare correct and thatsolace.oauth.client.token.fetch.timeoutis long enough for the authorization server to respond. -
For file mode, confirm that the file at
solace.oauth.client.access-tokencontains a currently valid token and that another process rotates it before it expires. A token has already lost some of its lifetime by the time the connector reads it from the file, so keep the gap between writing the file and the connector reading it small. If the rotating process doesn't update the file in time, the connector retries perreconnectRetriesand then stops if authentication doesn't succeed. -
If the authorization server uses a certificate that your JVM does not trust by default, configure
solace.oauth.client.auth-server.truststore.fileandtruststore.password. -
If authentication fails only after the job has been running successfully for a while, the token likely expired before
solace.oauth.client.token.refresh.intervalfired. Reduce the interval. -
If the error message is the literal text
null, the authorization server returned a response the connector couldn't parse as an OAuth error (for example, an HTML error page). Check the authorization server's own logs for the real cause. On the 4.x line, this case instead reports the HTTP status and response body.
-
- Queue Not Found or Consume Errors on the Source Side
-
If the source fails to start consuming:
-
Confirm that the queue named in the
queueoption already exists on the event broker. This connector does not create the consume queue. -
Confirm that the client username has consume permission on that queue.
-
Confirm that the queue has the topic subscriptions you expect data to arrive on.
-
- Checkpoint or LVQ Errors
-
If checkpointing fails, or the job cannot resume from where it left off:
-
Confirm that the client username's ACL allows publish and subscribe access to the LVQ topic (
lvq.topic). -
If an event broker administrator pre-created the LVQ, confirm that:
-
The LVQ is Exclusive.
-
The LVQ has a Spool Quota of
0. -
The administrator assigned your client username as the LVQ's owner.
-
-
At every startup, the connector asks the event broker to create the LVQ. If the LVQ already exists, the event broker does not create it again and does not return an error. Because the connector always sends this request, your client profile must always allow clients to create endpoints, even if an administrator pre-created the LVQ. If this permission is not enabled, the connector fails to start with a permission error (subcode 36), whether or not the LVQ exists. Enable the permission as follows:
-
For an event broker service, enable Allow Client to Create Endpoints in the client profile settings. For more information, see Client Profile Settings.
-
For a software event broker or an appliance event broker, enable
allow-guaranteed-endpoint-createin the client profile's message spool configuration. For more information, see Allowing Clients to Create Guaranteed Endpoints.
-
-
Ensure that no other client deletes the LVQ or consumes or deletes its messages. Any of these actions prevents the connector from resuming from the last checkpoint.
-
If a restart fails with a runtime error stating that the incoming message and the checkpointed message belong to different replication groups and cannot be compared, the connector cannot determine whether it already processed the incoming message. To log an error and process the message without deduplication, instead of failing the query, set
ignoreCheckpointMessageIdComparisonErrortotrue. This option takes effect only ifackLastProcessedMessagesistrue, and the connector ignores it if you setreplayStrategy. -
On Databricks, if checkpoint writes to a Unity Catalog Volume fail, confirm that the service principal named by
databricksClientId/databricksHost/databricksClientSecretstill has access to both the Volume and the secret scope and that its secret hasn't expired mid-rotation. Both the rotation and the connector's next refresh (databricksClientSecretRefreshInterval) must complete before the old secret expires.
-
- Sink Is Not Publishing Messages
-
If rows are not reaching the event broker:
-
Confirm that either the
topicoption is set or every row has a non-nullTopiccolumn value. A row missing both raises an exception rather than falling back to a default. -
Confirm that every row has a non-null
Payloadcolumn and that it has thebyte[]type. The connector rejects other payload types, and a nullPayloadraises an exception. -
Confirm that every row has a non-null
Idcolumn value. The connector requires it for every row, and a missing value raises an exception. -
If you are using
forEachBatchwith client certificate or OAuth authentication, note that this write path currently still requiresusernameandpasswordto be set. UsewriteStreaminstead if this is a problem. -
If commits are timing out with a
SolacePublishAckTimeoutException, increasepublishAckTimeoutor setpublishAckTimeoutFailOnErrortofalseto log timeouts instead of failing the job.
-
- Cast, Arithmetic, or Division Errors After Upgrading to the 4.x Line
-
Apache Spark 4.0 turns on ANSI SQL mode by default, so invalid casts, arithmetic overflow, and division by zero now raise errors instead of returning
NULL. This is a Apache Spark platform change, not a connector change. Setspark.sql.ansi.enabled=falseto restore the previous behavior or fix the underlying data issue. See Migrating to the 4.x Line. - Slow Throughput
-
If source throughput is lower than expected:
-
Increase
partitionsto add concurrent consumers, up to one per available Spark executor core. Settingpartitionsto0to auto-scale one consumer per worker node can create an unbounded number of consumers, so use it only if a fixedpartitionsvalue isn't sufficient. -
Set the queue's Maximum Delivered Unacknowledged Messages per Flow to roughly twice
batchSizeor equal tobatchSizeif messages are large. -
If the queue often has fewer messages available than
batchSize, reducequeue.receiveWaitTimeoutso the connector doesn't wait the full timeout to fill a batch. -
Drop the
Idcolumn before writing to your output sink if you don't need to replay from a specific message later. In benchmarking, dropping this column alone cut Spark per-batch write (addBatch) time from about 3 seconds to 1-2 seconds and raised throughput by roughly 1.5×. See Performance Considerations. -
If your rows include other large or unnecessary columns before writing to a downstream sink, consider dropping them too. Extra columns increase per-batch processing time.
-
- Messages Processed Out of Order
-
With
partitionsgreater than1, consumers compete for messages on the same queue, and the Spark task scheduling, not queue arrival order, determines processing order. If strict ordering is required, either handle out-of-order data downstream or use a single consumer (partitionsset to1) per queue and distribute load across multiple queues using an appropriate topic hierarchy. - Duplicate Messages
-
The
ackLastProcessedMessagesoption detects duplicates by comparing against the last processed checkpoint only. The connector has no visibility into whether the downstream system already processed a message. Duplicates can also result from the event broker redelivering uncommitted batches, an idle-timeout connection close with messages still pending acknowledgment, Spark checkpoint write failures during an instance crash, or cluster reinitialization. In any of these cases, add deduplication at the downstream system rather than relying onackLastProcessedMessagesalone.