Troubleshooting

The following troubleshooting tips might help you resolve issues with the Solace Connector for Apache Spark. If problems persist, contact Solace.

Batch Read Returns No Data or Fails

Only streaming reads are supported. If you call spark.read.format("solace") instead of spark.readStream.format("solace"), the read fails. Use readStream for the source side.

SolaceInvalidPropertyException at Startup
  • The exception message "Please provide Solace Queue in configuration options" means the queue option is missing. Add it before starting the source.

  • The exception message "Please set batch size greater than zero" means the batchSize parameter was set to a negative value. The connector rejects a negative value rather than treating it as unlimited; use 0 for no per-partition limit or a positive integer to cap how many messages each partition reads per micro-batch.

NoSuchMethodError, ClassNotFoundException, or Similar Errors at Startup

These typically indicate a version mismatch between the connector, Spark, and Scala, or a stale JAR file left on the classpath:

  • Confirm that the connector version matches your Spark/Scala/Java baseline. See Supported Apache Spark Releases.

  • Remove any earlier connector JAR file from the classpath before installing a new version, and restart your Spark environment. This is especially important if you move across the 3.x/4.x boundary, because a leftover 3.x (Scala 2.12) JAR file is binary-incompatible with the 4.x line's Scala 2.13 runtime and causes exactly this kind of error.

  • On Databricks, a Maven library install failing outright (rather than the job failing at runtime) on a Standard or Shared access-mode cluster usually means the coordinate isn't on the workspace's Unity Catalog allowlist. Ask a metastore admin to add it, or use a Dedicated access-mode cluster instead.

Connection Errors

If the job fails to connect to the event broker:

  • Verify that host, vpn, username, and password (or the equivalent client certificate or OAuth options) are correct.

  • Confirm that the client username's ACL permits connecting to the specified Message VPN.

  • Increase connectRetries and reconnectRetries if the event broker is temporarily unavailable during job startup.

OAuth Authentication Failures

If OAuth authentication fails:

  • Confirm that solace.apiProperties.AUTHENTICATION_SCHEME is set to AUTHENTICATION_SCHEME_OAUTH2.

  • For server-fetch mode, verify that solace.oauth.client.auth-server-url, client-id, and credentials.client-secret are correct and that solace.oauth.client.token.fetch.timeout is long enough for the authorization server to respond.

  • For file mode, confirm that the file at solace.oauth.client.access-token contains a currently valid token and that another process rotates it before it expires. A token has already lost some of its lifetime by the time the connector reads it from the file, so keep the gap between writing the file and the connector reading it small. If the rotating process doesn't update the file in time, the connector retries per reconnectRetries and then stops if authentication doesn't succeed.

  • If the authorization server uses a certificate that your JVM does not trust by default, configure solace.oauth.client.auth-server.truststore.file and truststore.password.

  • If authentication fails only after the job has been running successfully for a while, the token likely expired before solace.oauth.client.token.refresh.interval fired. Reduce the interval.

  • If the error message is the literal text null, the authorization server returned a response the connector couldn't parse as an OAuth error (for example, an HTML error page). Check the authorization server's own logs for the real cause. On the 4.x line, this case instead reports the HTTP status and response body.

Queue Not Found or Consume Errors on the Source Side

If the source fails to start consuming:

  • Confirm that the queue named in the queue option already exists on the event broker. This connector does not create the consume queue.

  • Confirm that the client username has consume permission on that queue.

  • Confirm that the queue has the topic subscriptions you expect data to arrive on.

Checkpoint or LVQ Errors

If checkpointing fails, or the job cannot resume from where it left off:

  • Confirm that the client username's ACL allows publish and subscribe access to the LVQ topic (lvq.topic).

  • If an event broker administrator pre-created the LVQ, confirm that:

    • The LVQ is Exclusive.

    • The LVQ has a Spool Quota of 0.

    • The administrator assigned your client username as the LVQ's owner.

  • At every startup, the connector asks the event broker to create the LVQ. If the LVQ already exists, the event broker does not create it again and does not return an error. Because the connector always sends this request, your client profile must always allow clients to create endpoints, even if an administrator pre-created the LVQ. If this permission is not enabled, the connector fails to start with a permission error (subcode 36), whether or not the LVQ exists. Enable the permission as follows:

    • For an event broker service, enable Allow Client to Create Endpoints in the client profile settings. For more information, see Client Profile Settings.

    • For a software event broker or an appliance event broker, enable allow-guaranteed-endpoint-create in the client profile's message spool configuration. For more information, see Allowing Clients to Create Guaranteed Endpoints.

  • Ensure that no other client deletes the LVQ or consumes or deletes its messages. Any of these actions prevents the connector from resuming from the last checkpoint.

  • If a restart fails with a runtime error stating that the incoming message and the checkpointed message belong to different replication groups and cannot be compared, the connector cannot determine whether it already processed the incoming message. To log an error and process the message without deduplication, instead of failing the query, set ignoreCheckpointMessageIdComparisonError to true. This option takes effect only if ackLastProcessedMessages is true, and the connector ignores it if you set replayStrategy.

  • On Databricks, if checkpoint writes to a Unity Catalog Volume fail, confirm that the service principal named by databricksClientId/databricksHost/databricksClientSecret still has access to both the Volume and the secret scope and that its secret hasn't expired mid-rotation. Both the rotation and the connector's next refresh (databricksClientSecretRefreshInterval) must complete before the old secret expires.

Sink Is Not Publishing Messages

If rows are not reaching the event broker:

  • Confirm that either the topic option is set or every row has a non-null Topic column value. A row missing both raises an exception rather than falling back to a default.

  • Confirm that every row has a non-null Payload column and that it has the byte[] type. The connector rejects other payload types, and a null Payload raises an exception.

  • Confirm that every row has a non-null Id column value. The connector requires it for every row, and a missing value raises an exception.

  • If you are using forEachBatch with client certificate or OAuth authentication, note that this write path currently still requires username and password to be set. Use writeStream instead if this is a problem.

  • If commits are timing out with a SolacePublishAckTimeoutException, increase publishAckTimeout or set publishAckTimeoutFailOnError to false to log timeouts instead of failing the job.

Cast, Arithmetic, or Division Errors After Upgrading to the 4.x Line

Apache Spark 4.0 turns on ANSI SQL mode by default, so invalid casts, arithmetic overflow, and division by zero now raise errors instead of returning NULL. This is a Apache Spark platform change, not a connector change. Set spark.sql.ansi.enabled=false to restore the previous behavior or fix the underlying data issue. See Migrating to the 4.x Line.

Slow Throughput

If source throughput is lower than expected:

  • Increase partitions to add concurrent consumers, up to one per available Spark executor core. Setting partitions to 0 to auto-scale one consumer per worker node can create an unbounded number of consumers, so use it only if a fixed partitions value isn't sufficient.

  • Set the queue's Maximum Delivered Unacknowledged Messages per Flow to roughly twice batchSize or equal to batchSize if messages are large.

  • If the queue often has fewer messages available than batchSize, reduce queue.receiveWaitTimeout so the connector doesn't wait the full timeout to fill a batch.

  • Drop the Id column before writing to your output sink if you don't need to replay from a specific message later. In benchmarking, dropping this column alone cut Spark per-batch write (addBatch) time from about 3 seconds to 1-2 seconds and raised throughput by roughly 1.5×. See Performance Considerations.

  • If your rows include other large or unnecessary columns before writing to a downstream sink, consider dropping them too. Extra columns increase per-batch processing time.

Messages Processed Out of Order

With partitions greater than 1, consumers compete for messages on the same queue, and the Spark task scheduling, not queue arrival order, determines processing order. If strict ordering is required, either handle out-of-order data downstream or use a single consumer (partitions set to 1) per queue and distribute load across multiple queues using an appropriate topic hierarchy.

Duplicate Messages

The ackLastProcessedMessages option detects duplicates by comparing against the last processed checkpoint only. The connector has no visibility into whether the downstream system already processed a message. Duplicates can also result from the event broker redelivering uncommitted batches, an idle-timeout connection close with messages still pending acknowledgment, Spark checkpoint write failures during an instance crash, or cluster reinitialization. In any of these cases, add deduplication at the downstream system rather than relying on ackLastProcessedMessages alone.