Nov 3, 2020

Streaming Geopoint data from Kafka to Elasticsearch

Streaming data from Kafka to Elasticsearch is easy with Kafka Connect - you can see how in this tutorial and video.

One of the things that sometimes causes issues though is how to get location data correctly indexed into Elasticsearch as geo_point fields to enable all that lovely location analysis. Unlike data types like dates and numerics, Elasticsearch’s Dynamic Field Mapping won’t automagically pick up geo_point data, and so you have to do two things:

Oct 7, 2020

ksqlDB - How to model a variable number of fields in a nested value (`STRUCT`)

There was a good question on StackOverflow recently in which someone was struggling to find the appropriate ksqlDB DDL to model a source topic in which there was a variable number of fields in a STRUCT.

Oct 5, 2020

Streaming XML messages from IBM MQ into Kafka into MongoDB

Let’s imagine we have XML data on a queue in IBM MQ, and we want to ingest it into Kafka to then use downstream, perhaps in an application or maybe to stream to a NoSQL store like MongoDB.

Note	This same pattern for ingesting XML will work with other connectors such as JMS and ActiveMQ.

Oct 1, 2020

Ingesting XML data into Kafka - Option 3: Kafka Connect FilePulse connector

👉 Ingesting XML data into Kafka - Introduction

We saw in the first post how to hack together an ingestion pipeline for XML into Kafka using a source such as curl piped through xq to wrangle the XML and stream it into Kafka using kafkacat, optionally using ksqlDB to apply and register a schema for it.

The second one showed the use of any Kafka Connect source connector plus the kafka-connect-transform-xml Single Message Transformation. Now we’re going to take a look at a source connector from the community that can also be used to ingest XML data into Kafka.

Oct 1, 2020

Ingesting XML data into Kafka - Option 2: Kafka Connect plus Single Message Transform

We previously looked at the background to getting XML into Kafka, and potentially how [not] to do it. Now let’s look at the proper way to build a streaming ingestion pipeline for XML into Kafka, using Kafka Connect.

If you’re unfamiliar with Kafka Connect, check out this quick intro to Kafka Connect here. Kafka Connect’s excellent plugable architecture means that we can pair any source connector to read XML from wherever we have it (for example, a flat file, or a MQ, or anywhere else), with a Single Message Transform to transform the XML into a payload with a schema, and finally a converter to serialise the data in a form that we would like to use such as Avro or Protobuf.

Oct 1, 2020

Ingesting XML data into Kafka - Option 1: The Dirty Hack

👉 Ingesting XML data into Kafka - Introduction

What would a blog post on rmoff.net be if it didn’t include the dirty hack option? 😁

The secret to dirty hacks is that they are often rather effective and when needs must, they can suffice. If you’re prototyping and need to JFDI, a dirty hack is just fine. If you’re looking for code to run in Production, then a dirty hack probably is not fine.

Oct 1, 2020

Ingesting XML data into Kafka - Introduction

XML has been around for 20+ years, and whilst other ways of serialising our data have gained popularity in more recent times (such as JSON, Avro, and Protobuf), XML is not going away soon. Part of that is down to technical reasons (clearly defined and documented schemas), and part of it is simply down to enterprise inertia - having adopted XML for systems in the last couple of decades, they’re not going to be changing now just for some short-term fad.

Oct 1, 2020

`abcde` - Error trying to calculate disc ids without lead-out information

Short & sweet to help out future Googlers. Trying to use abcde I got the error:

[WARNING] something went wrong while querying the CD... Maybe a DATA CD or the CD is not loaded?
[WARNING] Error trying to calculate disc ids without lead-out information.

Oct 1, 2020

IBM MQ on Docker - Channel was blocked

Running IBM MQ in a Docker container and the client connecting to it was throwing repeated Channel was blocked errors.

Sep 30, 2020

Setting key value when piping from jq to kafkacat

One of my favourite hacks for getting data into Kafka is using kafkacat and stdin, often from jq. You can see this in action with Wi-Fi data, IoT data, and data from a REST endpoint. This is fine for getting values into a Kafka message - but Kafka messages are key/value, and being able to specify a key is can often be important.

Here’s a way to do that, using a separator and some jq magic. Note that at the moment kafkacat only supports single byte separator characters, so you need to choose carefully. If you pick a separator that also appears in your data, it’s possibly going to have unintended consequences.

rmoff’s random ramblings

✨ Data Engineering, Kafka, and other random geekery 🤓

Streaming Geopoint data from Kafka to Elasticsearch

ksqlDB - How to model a variable number of fields in a nested value (`STRUCT`)

Streaming XML messages from IBM MQ into Kafka into MongoDB

Ingesting XML data into Kafka - Option 3: Kafka Connect FilePulse connector

Ingesting XML data into Kafka - Option 2: Kafka Connect plus Single Message Transform

Ingesting XML data into Kafka - Option 1: The Dirty Hack

Ingesting XML data into Kafka - Introduction

`abcde` - Error trying to calculate disc ids without lead-out information

IBM MQ on Docker - Channel was blocked

Setting key value when piping from jq to kafkacat