Comparing RESTful APIs and SOAP APIs Using MuleSoft as an Example

Simple Object Access Protocol (SOAP) and Representational State Transfer (REST) are two answers to the same question: how to access web services. The choice initially may seem easy, but at times it can be surprisingly difficult.

SOAP is a standards-based web services access protocol that has been around for a while and enjoys all of the benefits of long-term use. Originally developed by Microsoft, SOAP really isn’t as simple as the acronym would suggest. https://goo.gl/o9RsrF #DataIntegration #ML

Your Data Is Sound, But How’s Your Dashboard? 5 Aspects to Consider

One of the biggest problems in data management and data science is being able to obtain “good” data. You need to gather sufficient data from a substantial array of subjects who fit your study’s requirements, and ensure the accuracy of the data… otherwise, any conclusions you draw could be biased or skewed.

But assume for a moment that your data is already solid. That’s no guarantee of success, unfortunately: It’s like having all the ingredients of a pizza in one place but lacking the ability to tie those ingredients together, and cook them appropriately.

Without the latter, you may not get the final product you seek. In addition to considering the quality of your data, consider the quality of your dashboard; it’s more important than you might assume.

Why Your Dashboard Matters

Here are some of the reasons your dashboard should matter as much as your data.

* Access. First, you need to be able to call up as much of the data as possible. If your dashboard highlights a handful of key variables, but makes others harder to see or understand, it could lead you to false conclusions or undersell what you’ve been able to gather.
* Manipulation. Your dashboard is also what empowers you to tweak different variables, generate comparative reports, play around with different timeframes and demographics, and ultimately give you the “full picture” of your subject matter.
* Showcasing. Depending on your company and your position, you’ll probably need to make sure other people can see and understand the data before you can reap its true content. That’s where visualizations come into play. Your dashboard should make it easy for people to wrap their minds around your findings, regardless of whether they were involved in accumulating them.

Key Considerations

In the current data-driven marketplace,there are hundreds of unique dashboards  you can use to analyze and display information. How to choose which would be best for your needs?

* Ease of use. First, you should make sure your dashboard is easy to use, both for you and the others on your team. If it sucks up a few hours of study and playing around to learn the basic functions, it will probably include features you miss entirely. Beyond that, it may cost hours of company time to get new hires up to speed, and anyone outside your team who tries to use or view the platform could be baffled. Your dashboard should be more or less intuitive, if possible.
* Variable controls. You’ll also need a platform that has sufficient variable controls, which will allow you to create your own custom reports and change them dynamically as you spend more time on the platform. It should be relatively easy to account for new variables, reframe your data with new parameters, and dig deeper to unearth further insights. Cookie-cutter reports and controls aren’t likely to meet your needs in today’s business arena.
* Design aesthetics. Don’t discount the value of the aesthetics of your dashboard. Your data visuals should exist to tell a story about your data, both to people on your team and outside of it. If that story is hard to follow, or looks boring, your audience either won’t be able to draw accurate conclusions, or won’t be inspired to do so.
* Feature approachability. What good is a dashboard with a ton of features if you only need a few of them to obtain the results you need? You might be tempted to opt for a dashboard that offers lots of bells and whistles, but those perks won’t necessarily offer the best fit for your organization. Instead, find a platform with features that will contour to your needs, and are relatively easy to find and master.
* Access and share-ability. Finally, you need to make sure there won’t be any obstacles with regard to access or share-ability. Most organizations will want a dashboard with multiple “access” levels, including administrative and view-only accounts. You should also weigh how customized reports can be displayed, exported, and circulated to others. This is one of the most important functions of data gathering and distribution.

Your dashboard is more than just a user interface that allows you to get access to raw information. It’s a filter and a platform that can help you get the most out of your data.

Think carefully before you make the decision, and keep auditing that decision as you use the platform in your daily work, because something better may be on the way, or already available. https://goo.gl/UM89TP #DataScience #Cloud

The Basic Stuff of Machine Learning

By now anyone who reads virtually any trade magazine has been hearing incessantly about how machine learning is going to transform their industry in profound ways. Marketers will be able to read potential customers’ minds, farms will produce unprecedented yields, doctors will be able to stem diseases before they begin to form. And of course, we’ve all heard how machine learning will eventually take our jobs. It may very well be said of machine learning that there never have been so many wild predictions made about something which the majority of the public knows so little. So what exactly is machine learning? And what can we reasonably expect in the next ten years? And of course the question that has been plaguing us all: will the machines rise up and destroy us?

Will the real Pinocchio please stand?

Machine learning, at its essence, is a form of AI that involves the attempt to get computers to perform actions without being explicitly programmed to do so. It’s divided into two basic categories of algorithms. The first is supervised learning, which involves “training” the algorithm with data to develop the ability to recognize certain patterns within data, in order to categorize or predict. For example, if you wanted to write a program that recognized photos of Pinocchio, you might feed it a thousand pictures, 500 with Pinocchio, and 500 with other random people, indicating “Pinocchio”, or “Not Pinocchio” for each observation. With each observation, it would learn a little bit more, eventually being able to distinguish between Pinocchio and other long-nosed characters like Scrooge or Gonzo from The Muppets. This is similar to how you might train a toddler to recognize categories of things–the difference, of course, being that the algorithm can ingest training data much more quickly. We see the applications of this all over the place, from voice recognition to types of fraud detection that involve pattern recognition, with regression, Support Vector Machines, Bayes and decision trees being popular algorithm types.

A second category of machine learning is known as “unsupervised learning”. It’s a bit more esoteric, and it’s sometimes easier to understand by first explaining how it differs from supervised learning. If supervised learning is compared to explaining to someone how to navigate from LA to NYC, and then turning them loose and letting them make the trip, unsupervised learning is more like a Lewis and Clark expedition, where you send them out to map the landscape. Unsupervised learning algorithms such as k-means are often used to categorize data into groups according to similarities. When your 2-year old child goes rummaging through your kitchen, he or she is engaging in unsupervised learning as it makes sense of a host of shiny new objects.

Will the long-nosed talking wooden puppet please stand?

Another way to think about the difference between supervised and unsupervised learning is that in supervised learning, the you already know what the data is and you’re telling the application as much. Back to the Pinocchio example, you already know what our wooden friend looks like–you’ve labeled the data. In unsupervised learning, the data isn’t labeled–you’re turning the algorithm loose, in a sense, and letting it swim around in a dataset until it starts to make sense of it. So if you gave it 1000 pictures of characters–500 of them different representations of Pinocchio, and 500 of them other random characters, none of them who were the same–and you asked the algorithm to identify the character that was the same, the percentage of photos of “Pinocchio” photos correctly lumped in the same category would demonstrate the success of your algorithm.

From this starting point, machine learning starts to branch off into more complex, and very exciting territories which combine and build on various aspects of supervised and unsupervised learning. Deep learning, for instance, attempts to mimic the human brain by layering one abstraction on top of another. Think of, for example, how a child first learns to distinguish between an animal and a stuffed animal, and then later learns to distinguish between a dog and a cat, a mammal and a reptile, and on to more complex classifications (which eventually form the basis of more complex decisions, like whether or not to pet that dog that’s foaming at the mouth).   

Will the machines rise up and destroy us?

With the likes of Stephen Hawking, Elon Musk and Mark Zuckerberg heatedly debating this, I’m not going to offer any opinion on the matter. Maybe they will, or perhaps they won’t. But, it is informative to look at where we currently stand in our progress toward intelligent machines (which may or may not rule the world).

Arend Hintze of Michigan State came up with 4 classifications of AI, including ‘reactive machines’, where algorithms react to data they’re being fed but have no ability to contextualize; ‘limited memory’, where past experiences begin to inform decisions; ‘theory of mind’ where computers start to recognize that others have beliefs, desires and intentions; and ‘self-awareness’ where computers become sentient beings–like C3PO, HAL or, heaven forbid, Skynet.

So where are we today? As far as we know, we’re still in the “Limited Memory” stage. How long will it be until we cross the next chasm? It’s impossible to say, but we do see things progressing fast. Arguably, the Turing Test was passed a few years ago by the Eugene Goostman program a few years ago (though this claim has been heatedly contested). What is incontestable, however, is that corporations and governments are pouring billions of dollars into research on how this technology can be applied to just about every sector of business and life. Clearly, we are about to see some pretty amazing things from machine learning. Anticipate that before too long, machine learning algorithms will not just be able to recognize Pinocchio, but also to spot whether or not he’s lying even before his nose grows! https://goo.gl/rx6VJj #DataScience #Cloud