Posts

Showing posts with the label hadoop

How-to: Index Scanned PDFs at Scale Using Fewer Than 50 Lines of Code

Image
Learn how to use OCR tools, Apache Spark, and other Apache Hadoop components to process PDF images at scale. Optical character recognition (OCR) technologies have advanced significantly over the last 20 years. However, during that time, there has been little or no effort to marry OCR with distributed architectures such as Apache Hadoop to process large numbers of images in near-real time. In this post, you will learn how to use standard open source tools along with Hadoop components such as Apache Spark, Apache Solr, and Apache HBase to do just that for a medical device information use case. Specifically, you will use a public  dataset  to convert narrative text into searchable fields. Although this example concentrates on medical device information, it can be applied in many other scenarios where processing and persisting images is required. Insurance companies, for example, can make all their scanned documents in claims files searchable for better claim resolu...

Apache Sentry is Now a Top-Level Project

Image
Source: Cloudera The following post was originally published by the Sentry community at apache.org. We re-publish it here for your convenience. We are very excited to announce that  Apache Sentry  has graduated out of Incubator and is now an Apache Top-Level Project! Sentry, which provides centralized fine-grained access control on metadata and data stored in Apache Hadoop clusters, was introduced as an Apache Incubator project back in August 2013. In the past two and a half years, the development community grew significantly to a large number of contributors from various organizations. Upon graduation, there were more than 50 contributors, 31 of whom had become committers. What’s Sentry? While Hadoop has strong security at the filesystem level, it lacked the granular support needed to adequately secure access to data by users and BI applications. This problem forces users to make a choice: either leave data unprotected or lock out users entirely. Most of the time, the...

Meet Cloudera’s Apache Spark Committers

Image
From Cloudera Blog, thanks to    Justin Kestelyn   for valuable post in cloudera blog. The super-active Apache Spark community is exerting a strong gravitational pull within the Apache Hadoop ecosystem. I recently had that opportunity to ask Cloudera’s Apache Spark committers (Sean Owen, Imran Rashid [PMC], Sandy Ryza, and Marcelo Vanzin) for their perspectives about how the Spark community has worked and is working together, and the work to be done via the  One Platform initiative  to make the Spark stack enterprise-ready. Recently, Apache Spark has become the most currently active project in the Apache Hadoop ecosystem (measured by number of contributors/commits over time), if not the entire ASF. Why do you think that is? Owen: Partly because of scope: Apache Spark has been many sub-projects under an umbrella from the start, some large and complex in their own right, and has tacked on several more in just the last six months. Culture is anothe...

Working with Apache Spark: Or, How I Learned to Stop Worrying and Love the Shuffle

this post is from cloudera blog, thanks to IIya Ganelin Our thanks to Ilya Ganelin, Senior Data Engineer at Capital One Labs, for the guest post below about his hard-earned lessons from using Spark. I started using Apache Spark in late 2014, learning it at the same time as I learned Scala, so I had to wrap my head around the various complexities of a new language as well as a new computational framework. This process was a great in-depth introduction to the world of Big Data (I previously worked as an electrical engineer for Boeing), and I very quickly found myself deep in the guts of Spark. The hands-on experience paid off; I now feel extremely comfortable with Spark as my go-to tool for a wide variety of data analytics tasks, but my journey here was no cakewalk. Capital One’s original use case for Spark was to surface product recommendations for a set of 25 million users and 10 million products, one of the largest datasets available for this type of modeling. Moreover, we had ...

Security, Hive-on-Spark, and Other Improvements in Apache Hive 1.2.0

Apache Hive 1.2.0, although not a major release, contains significant improvements. Recently, the Apache Hive community moved to a more frequent, incremental release schedule. So, a little while ago, we  covered the Apache Hive 1.0.0 release  and explained how it was renamed from 0.14.1 with only minor feature additions since 0.14.0. Shortly thereafter,  Apache Hive 1.1.0  was released (renamed from Apache Hive 0.15.0), which included more significant features—including  Hive-on-Spark . Last week, the community released  Apache Hive 1.2.0 . Although a more narrow release than Hive 1.1.0, it nevertheless contains improvements in the following areas: New Functionality Support for Apache Spark 1.3 ( HIVE-9726 ), enabling dynamic executor allocation and impersonation Support for integration of Hive-on-Spark with Apache HBase ( HIVE-10073 ) Support for numeric partition columns with literals ( HIVE-10313 ,  HIVE-10307 ) Support for Union Dist...

Apache Hadoop Infrastructure Considerations and Best Practices

Image
Thanks to Lisa Sensmeier and hortonworks Link Bit Refinery is a Hortonworks Technical Partner and recently certified with HDP. Bit Refinery is a VMware© Cloud Infrastructure-as-a-Service (IaaS) provider featuring virtualization technology hosted within their fully redundant virtual data centers. Bit Refinery offers a  hosted Hortonworks Sandbox  providing an easy way to experience and learn Hadoop with ease. All the tutorials available from the Hortonworks Sandbox work just as if you were running a localized version of the Sandbox. Brandon Hieb, Managing Partner at Bit Refinery, is our guest blogger, and in this blog, he provides insight to virtualizing Hadoop infrastructures. Here at Bit Refinery we provide infrastructure for companies large and small which includes a variety of big data applications running on both bare-metal and VMware servers. With this new technology constantly changing, it’s hard to keep up with the different required resources needed which co...

Hortonworks: A New Way for Certification

Image
for more info/Source : Click Here Hortonworks is excited to announce that our first hands-on, performance based certification exam is now available! The HDP Certified Developer (HDPCD) exam is designed for Hadoop developers working with frameworks like  Pig ,  Hive ,  Sqoop  and  Flume . This new approach to Hadoop certification is designed to allow individuals an opportunity to prove their Hadoop skills in a way that is recognized in the industry as meaningful and relevant to on-the-job performance. Instead of multiple-choice questions, the exam consists of tasks executed on a live, three-node Hortonworks Data Platform cluster: Exam Objectives The exam has three main categories of tasks that involve: Data ingestion Data transformation Data analysis Click  here  to view a detailed list of objectives for the HDPCD exam.

Using Apache Sqoop for Load Testing

Reference: Cloudera Our thanks to Montrial Harrell, Enterprise Architect for the State of Indiana, for the guest post below. Recently, the State of Indiana has begun to focus on how enterprise data management can help our state’s government operate more efficiently and improve the lives of our residents. With that goal in mind, I began this journey just like everyone else I know: with an interest in learning more about Apache Hadoop. I started learning Hadoop via a virtual server onto which I installed CDH and worked through a few online tutorials. Then, I learned a little more by reading blogs and documentation, and by trial and error. Eventually, I decided to experiment with a classic Hadoop use case: extract, load, and transfer (ELT). In most cases, ELT allows you to offload some resource-intensive data transforms in favor of Hadoop’s MPP-like functionality, thereby cutting resource usage on the current ETL server at a relatively low cost. This functionality is in part del...

New in CDH 5.3: Transparent Encryption in HDFS

Image
Thanks to Cloudera for article and updated Version of CDH 5.3 Support for transparent, end-to-end encryption in HDFS is now available and production-ready (and shipping inside  CDH 5.3  and later). Here’s how it works. Apache Hadoop 2.6 adds support for transparent encryption to HDFS. Once configured, data read from and written to specified HDFS directories will be transparently encrypted and decrypted, without requiring any changes to user application code. This encryption is also end-to-end, meaning that data can only be encrypted and decrypted by the client. HDFS itself never handles unencrypted data or data encryption keys. All these characteristics improve security, and HDFS encryption can be an important part of an organization-wide data protection story. Cloudera’s HDFS and Cloudera Navigator Key Trustee (formerly Gazzang zTrustee) engineering teams did this work under  HDFS-6134  in collaboration with engineers at Intel as an extension of earlier...

The Early Release Books Keep Coming: This Time, Hadoop Security

Image
source: Thanks to CLOUDERA for technical updates Hadoop Security  is the latest book from Cloudera engineers in the  Hadoop ecosystem books canon . We are thrilled to announce the availability of the early release of  Hadoop Security , a new book about security in the Apache Hadoop ecosystem published by O’Reilly Media. The early release contains two chapters on System Architecture and Securing Data Ingest and is available in  O’Reilly’s catalog  and in  Safari Books . The goal of the book is to serve the experienced security architect that has been tasked with integrating Hadoop into a larger enterprise security context. System and application administrators also benefit from a thorough treatment of the risks inherent in deploying Hadoop in production and the associated how and why of Hadoop security. As Hadoop continues to mature and become ever more widely adopted, material must become specialized for the security architects tasked with ensur...

Getting Started with Big Data Architecture

Image
What does a “Big Data engineer” do, and what does “Big Data architecture” look like? In this post, you’ll get answers to both questions. Apache Hadoop has come a long way in its relatively short lifespan. From its beginnings as a reliable storage pool with integrated batch processing using the scalable, parallelizable (though inherently sequential) MapReduce framework, we have witnessed the recent additions of real-time (interactive) components like Impala for interactive SQL queries and integration with Apache Solr as a search engine for free-form text exploration. Getting started is now also a lot easier: Just install CDH, and all the Hadoop ecosystem components are at your disposal. But after installation, where do you go from there? What is a good first use case? How do you ask those “bigger questions”? Having worked with more customers running Hadoop in production than any other vendor, Cloudera’s field technical services team has seen more than its fair share of these use...

The New Hadoop Application Architectures Book is Here!

Image
Thanks to Cloudera(Source)                                 ---         Get this Copy I Every time follow Cloudera blog and updating their information to share for all technocrats There’s an important new addition coming to the Apache Hadoop book ecosystem. It’s now in early release! We are very happy to announce that the new Apache Hadoop book we have been writing for O’Reilly Media,  Hadoop Application Architectures , is now available as an early release! It contains the first two chapters and can be found in O’Reilly’s Catalog  and via  Safari .         The goal of this book is to give developers and architects guidance on architecting end-to-end solutions using Hadoop and tools in the ecosystem. We have split the book into two broad sections: the first section discusses various considerations for designing applications, and the second sect...

How-to: Create an IntelliJ IDEA Project for Apache Hadoop

Image
Thanks to  Charles Lamb, Source: Cloudera Link: here Prefer  IntelliJ IDEA over Eclipse? We’ve got you covered: learn how to get ready to contribute to Apache Hadoop via an IntelliJ project. It’s generally useful to have an IDE at your disposal when you’re developing and debugging code. When I first started working on HDFS,  I used Eclipse , but I’ve recently switched to JetBrains’  IntelliJ IDEA  (specifically, version 13.1 Community Edition). My main motivation was the ease of project setup in the face of Maven and  Google Protocol Buffers  (used in HDFS). The latter is an issue because the code generated by  protoc  ends up in one of the target subdirectories, which can be a configuration headache. The problem is not that Eclipse can’t handle getting these files into the classpath — it’s that in my personal experience, configuration is cumbersome and it takes more time to set up a new project whenever I make a new clone. Conversely...

Introducing : Cloudera LIVE

Image
Source: cloudera , Thanks to Cloudera Cloudera Live  is a new way to get started with Apache Hadoop, online. No downloads, no installations, no waiting. Watch tutorial videos and work with real-world examples of the complete Hadoop stack included with CDH, Cloudera’s completely open source Hadoop platform, to: Learn Hue, the Hadoop User Interface developed by Cloudera Query data using popular projects like Apache Hive, Apache Pig, Impala, Apache Solr, and Apache Spark (new!) Develop workflows using Apache Oozie

Experts Video - Hadoop World

Image
Watch the experts Doug Cutting, Cloudera's Chief Architect, and Todd Papaioannou  Splunk 's CTO as they share their thoughts on big data in Asia on  Bloomberg Television 's Singapore Sessions

Using Apache Hadoop and Impala with MySQL for Data Analysis

source: Cloudera blog;  Thanks to Alexander Rubin of Percona Apache Hadoop is commonly used for data analysis. It is fast for data loads and scalable. In a previous post I showed how to  integrate MySQL with Hadoop . In this post I will show how to export a table from  MySQL to Hadoop, load the data to  Cloudera Impala  (columnar format), and run reporting on top of that. For the examples below, I will use the “ontime flight performance” data from my  previous post . I’ve used  Cloudera Manager  to install Hadoop and Impala. For this test I’ve (intentionally) used an old hardware (servers from 2006) to show that Hadoop can utilize the old hardware and still scale. The test cluster consists of 6 datanodes. Below are the specs: Purpose Server specs Namenode, Hive metastore, etc + Datanodes 2x PowerEdge 2950, 2x L5335 CPU @ 2.00GHz, 8 cores, 16GB RAM, RAID 10 with 8 SAS drives Datanodes only 4x PowerEdge SC1425, 2x Xeon ...

What is a Big Data cluster?

Image
Very often I get the query `What is a cluster?` when discussing about Hadoop and Big Data. To keep it simple ` A cluster is a group or a network of machines wired together acting a single entity to work on a task which when run on a single machine takes much more longer time. ` The given task is split and processed by multiple machines in parallel and so that the task gets completed faster. Jesse Johnson puts it in simple and clear terms what a cluster is all about and how to design distributed algorithms here .                                                 In a Big Data cluster, the machines (or nodes) are neither as powerful as a server grade machine nor as dumb as a desktop machine. Having multiple (like in thousands) server grade ma...

Caching Proxy - Installation and Configuration

Image
Setting up a Hadoop cluster is all easy with a bit of familiarity with system and network administration. It's all interesting, the only frustrating thing is the downloading of the patches after the installation of the OS and the downloading of the packages for the softwares on top of OS. The downloads can go to all the way close to a GB also, which might take a couple of minutes to hours based on the internet bandwidth. Here is where caching tools really help. They will cache the downloaded packages to one of the designated local machine (lets call it the cache server) and the other machines can point to the cache server to get the packages. This way the packages are downloaded from the internet for the first time and from then on the local cache server will be used for getting the packages. This approach will not only save the network bandwidth, but will also make the whole installation process faster. For debian systems, apt-cacher-ng is designed to cache the ...

HIVE Installation & Setup Guide

Image
Pre-requisites Ubuntu / CentOS Hadoop 1.x/ 2.x , I prefer to install with 2.x Step –> 1: Download and Install Download the Hive from the Apache Download Mirror and i place it in /home/bigdata/Installations/ d irectory. $ cd /home/bigdata/Installation $ wget http://redrockdigimark.com/apachemirror/hive/stable/apache-hive-1.2.1-bin.tar.gz  ( i preferred to download hive-1.2.1 .tar.gz, as it is stable version) $ sudo tar xzf hive-1.2.1.tar.gz Step –> 2: After downloading and installation. Now we are moving to edit hive-env.sh file for Configuration. To configure hive, there I have installed and give permission to bigdata. In $HIVE_HOME/conf/hive-env.sh export JAVA_HOME=/opt/jdk1.80_10  Step –> 3: add hbase path to bashrc $ gedit .bashrc and add following lines to it #HIVE export HIVE_HOME=/home/bigdata/Installations/hive-1.2.1/ export PATH=$PATH:$HIVE_HOME/bin Step –> 4: Restart the terminal and start hadoop, th...

Hadoop Research Tips

Those who are interested to work on Hadoop, One commonly asked question that I got from these people  is  what Hadoop feature can I work on? Here are some items that I have in mind that are good topics for students to attempt if they want to work in Hadoop. Ability to make Hadoop scheduler resource aware, especially CPU, memory and IO resources. The current implementation is based on statically configured  slots. Abilty to make a map-reduce job take new input splits even after a map-reduce job has already started. Ability to dynamically increase replicas of data in HDFS based on access patterns. This is needed to handle hot-spots of data. Ability to extend the map-reduce framework to be able to process data that resides partly in memory. One assumption of the current implementation is that the map-reduce framework is used to scan data that resides on disk devices. But memory on commodity machines is becoming larger and larger. A cluster of 3000 machines with ...

Big Data Trendz