Apache Hadoop Admin Tips and Tricks

In this post I will share some tips I learned after using the Apache Hadoop environment for some years, and  doing many many workshops and courses. The information here considers Apache Hadoop around version 2.9, but it could definably be extended to other similar versions.

These are considerations for when building or using a Hadoop cluster. Some are considerations over the Cloudera distribution. Anyway, hope it helps! 

* Don’t use Hadoop for millions of small files. It overloads the namenode and makes it slower. It is not difficult to overload the namenode. Always check capability vs number of files. Files on Hadoop usually should be more than 100 MB.

* You have to have a 1 GB of memory for around 1 million files in the namenode.

* Nodes usually fail after 5 years. Node failures is one of the most frequent problems in Hadoop. Big companies like facebook and google should have node failures by the minute.

* The MySQL on Cloudera Manager does not have redundancy. This could be a point of failure.

* Information: the merging of fsimage files happens on the secondary namenode.

* Hadoop can cache blocks to improve performance. By default it caches 0. 

* You can set a parameter that sends an acknowledgment message from datanodes back to the namenode after only the first or second data block has been copied to the datanodes. That might make writing data  faster. 

* Hadoop has rack awareness: it knows which node is connected to witch switch. Actually, it it the Hadoop Admin who configures that.

* Files are checked from time to time to verify if there was any data corruption (usually every three weeks). This is possible because datanodes store files checksum.

* Log file stores by default 7 days.

* part-m-000 are from mapper and part-r-000 are from reducer jobs. The number in the end corresponds to the number of reducers that ran for that job. So part-r008 had 9 reducers  (starts from 0).

* You can change the log.level of mapper and reducers tasks yo get more information.

* mapreduce.reduce.log.level=DEBUG

* yarn server checks what spark did. localhost:4040 also shows what has been done.

* It is important to check where to put the namenode fsimage file.  You might want to replicate this file.

* You have to save a lot of disk space (25%) to dfs.datanode.du.reserve, for the shuffle phase.

* This phase is going to be written in disk, so there needs to be space!

* When you remove files, they stay on the .Trash directory after removing for a while. The default time is 1 day.

* You can build a lamdba architecture with flume. You can also specify if you want to put data in memory or disk flume.

* Regarding hardware, worker nodes need more cores for more processing. The master nodes don’t process that much.

* For the namenode you want more quality disks and better hardware (like raid – and raid makes no sense on worker nodes).

* The rule of thumb is: if you want to store 1 TB of data you have to have 4 TB space.

* Hadoop applications are typically not cpu bound. 

* Virtualization might give you some benefits (easier to manage), but it hits performance. Usually it brings between 5% and 30% of overhead.

* Hadoop does not support ipv6. You can disable ipv6. You can also disable selinux inside the cluster. Both give overhead.

* A good size for a starting cluster is around 6 nodes.
* Sometimes, when the clusters is too full, you might have to remove a small file to remove a bigger file.

That is it for now. I will try to write a part 2 soon. Let me know if there is anything I missed here! https://goo.gl/Bt1CDn #DataScience #Cloud

Are YOU the Outlier?

AI and machine learning are everywhere. Most decisions affecting every aspect of our lives are being made based on anomalies, classifications, and predictions. Even governmental decisions such as where will new schools be built may consider an enormous amount of demographic, geographic, and socioeconomic data to determine exactly which land will house the school – and developers are using similar data to buy up the plots they think the governments will choose.

We’re in an interesting position here at Binah. We’re creating and building the tools that help companies get a complete and true picture of their data – beyond anomalies, predictions, and classifications – to drive better models. Binah clients can do everything from better managing foreign currency exchange rates to maximize profits, dramatically increase customer retention and predict customer churn to analyzing IoT data that can send early warnings to your doctor that you might be at high risk for a heart attack.

Binah’s technology can increase business efficiency by better eliminating the outliers that may skew data, such as removing the unusual loan defaults so they can create better rules for offering loans to a greater majority of people.

The benefits of these machine learning and AI-based algorithmic models are clear. The school will be built in the area that will benefit the most children who will become school aged over the next 15 years. Loan policies will be changed to increase the possibilities of more people benefiting from home ownership and banks being repaid.
However, we’re putting all our trust in the numbers. All these decisions are based on probabilities – and life isn’t as certain as that.

Take insurance, for example. You are completely healthy, workout three times a week, and have low cholesterol and a reasonable BMI for your age. You show no signs of heart disease whatsoever. However, you had to give your medical history when applying, and your father had a heart attack and bypass surgery while your mother has high cholesterol. When the insurance company’s algorithmic model is applied to you, you become a very high risk for heart disease, and your rates are adjusted accordingly.

Those algorithms got you a fantastic low-interest car loan – even if most of your savings are eaten by those enormous health insurance premiums.

Where are you the “norm” and where are you the “outlier”? How do you prove it, in what situation? The more we depend on machine learning and big data as a society, it may become harder to fight injustice. Of course, it isn’t so easy to fight injustice today…

Maybe our dependence on big data, machine learning, and AI will at least help some of us in some parts of our lives – and punish us for being the outlier – or even being related to the outlier – in others.
David Maman is CEO, CTO & Co-founder of Binah.ai, whose out-of- the-box data science solutions leverage signal processing combined with machine learning and AI to create better models and accelerate delivery of the right answers to critical business questions. www.binah.ai https://goo.gl/cQBNBw #DataScience #Cloud