Bringing Discovery Within Data API Marketplaces out Into the Open

#ICYDK: I spend time reviewing each wave of data API marketplaces as they emerge on the landscape every couple of years. There are a number of reasons why these data marketplaces exist, ranging from supporting government agencies, NGOs, or for commercial purposes. One of the most common elements of API-driven data marketplaces that frustrates me is when they don’t do the hard work to expose the metadata around the databases, datasets, spreadsheets, and the raw data they are providing access to — making it very difficult to actually discover anything of interest.

You can see a couple examples of this with mLab, World Health Organization, Data.World, and others. While these platforms provide (sometimes) impressive ability to manage data stores, but they don’t always do a good job exposing the metadata of their catalogs as part of the available APIs. Dynamically generating API endpoints, documentation, and other resources based upon the data that is being published to their platforms. Leaving developers to do the digging, and making the investment to understand what is available on a platform. https://goo.gl/Pfh3gY #DataIntegration #ML

Monitor: Manager for Apache Kafka Clusters

With the new open source Streams Messaging Manager, you now have deep visibility into active Kafka flows across many producers, consumers, brokers, and topics.

Key Analysis: The true value of this tool is the ability to follow any message as it transits through multiple systems, clouds and hops. It is very easy to lose track of a consumer or producer when you have dozens or more topics and thousands of messages a second. You add in hybrid cloud and other connected systems and frameworks like Spark, Flink, NiFi, SAM, and Hadoop. https://goo.gl/D6REQW #DataIntegration #ML