Posts

Showing posts with the label New in CDH

Cloudera Data Hub: Where Agility Meets Control

Image
Cloudera’s new Data Hub cloud service, powered by  Cloudera Data Platform , enables users to seamlessly migrate on-premises data management and analytics workloads to the cloud as well as implement new cloud workloads in pursuit of your cloud-first data management strategy. On August 22 nd ,  Cloudera demonstrated its Data Hub  service during a webinar highlighting key business benefits, use cases, and product capabilities. Below is a brief overview of the topics covered and some of the most frequently asked questions from attendees. What is Cloudera Data Hub? Cloudera Data Hub  is a powerful cloud service on Cloudera Data Platform (CDP) that makes it easier, safer, and faster to build modern, mission-critical, data-driven applications with enterprise security, governance, scale, and control. The cloud-native service is powered by a suite of integrated open source technologies that delivers the widest range of analytical workloads such as data marts and data...

HBase Performance CDH5 (HBase1) vs CDH6 (HBase2)

Image
Thanks to Cloudera Blog HBase Customers upgrading to CDH 6 from CDH 5, will also get an HBase upgrade moving from HBase1 to HBase2. Performance is an important aspect customers consider.   We measured performance of CDH 5 HBase1 vs CDH 6 HBase2 using YCSB workloads to understand the performance implications of the upgrade on customers doing in-place upgrades (no changes to hardware).  About YCSB For our testing we used the  Yahoo! Cloud Serving Benchmark  (YCSB). YCSB is an open-source specification and program suite for evaluating retrieval and maintenance capabilities of computer programs. It is often used to compare relative performance of  NoSQL  database management systems. The original benchmark was developed by workers in the research division of  Yahoo!  who released it in 2010.  More info on YCSB at  https://github.com/brianfrankcooper/YCSB In our test environment YCSB @1TB data scale was used, and run workloads in...

New in CDH 5.4: Sensitive Data Redaction

Image
Thanks to  Michael Yoder The best data protection strategy is to remove sensitive information from everyplace it’s not needed Have you ever wondered what sort of “sensitive” information might wind up in Apache Hadoop log files? For example, if you’re storing credit card numbers inside HDFS, might they ever “leak” into a log file outside of HDFS? What about SQL queries? If you have a query like  select * from table where creditcard = '1234-5678-9012-3456' , where is that query information ultimately stored? This concern affects anyone managing a Hadoop cluster containing sensitive information. At Cloudera, we set out to address this problem through a new feature called  Sensitive Data Redaction , and it’s now available starting in Cloudera Manager 5.4.0 when operating on a CDH 5.4.0 cluster. Specifically, this feature addresses the “leakage” of sensitive information into channels unrelated to the flow of data–not the data stream itself. So, for example, Sensitive...

Big Data Trendz