Analytics

5 Strategies to Improve Secure Data Collaboration

Many organizations struggle to share data internally across departments and externally with partners, vendors, suppliers, and customers. They use manual methods such as emailing spreadsheets or executing batch processes that require extracting, copying, moving, and reloading data. These methods are notorious for their lack of stability and security, and most importantly, for the fact that by the time data is ready for consumption, it has often become stale.

A Cloud Data Platform for Data Science

Data scientists require massive amounts of data to build and train machine learning models. In the age of AI, fast and accurate access to data has become an important competitive differentiator, yet data management is commonly recognized as the most time-consuming aspect of the process. This white paper will help you identify the data requirements driving today's data science and ML initiatives and explain how you can satisfy those requirements with a cloud data platform that supports industry-leading tools.

Overview of the Operational Database performance in CDP

This article gives you an overview of Cloudera’s Operational Database (OpDB) performance optimization techniques. Cloudera’s Operational Database can support high-speed transactions of up to 185K/second per table and a high of 440K/second per table. On average, the recorded transaction speed is about 100K-300K/second per node. This article provides you an overview of how you can optimize your OpDB deployment in either Cloudera Data Platform (CDP) Public Cloud or Data Center.

Demand for Data Grows in Agriculture

Agriculture (Ag) is the oldest and largest industrial vertical in the world, and its importance continues to grow as it becomes more challenging for people to access healthy and fresh food. A recent Agriculture Analytics Market report, released by Markets and Markets, estimates that by 2023, the global agriculture analytics market size will grow from 585 million to 1.2 billion dollars as demands for real-time data analysis and improved operations increase.

Ask questions to BigQuery and get instant answers through Data QnA

Today, we’re announcing Data QnA, a natural language interface for analytics on BigQuery data, now in private alpha. Data QnA helps enable your business users to get answers to their analytical queries through natural language questions, without burdening business intelligence (BI) teams. This means that a business user like a sales manager can simply ask a question on their company’s dataset, and get results back that same way.

Powered by Fivetran Fuels Savvy Data Insights Platforms and Agencies

Powered by Fivetran (PBF) is a new offering for modern data insights platforms that provide analytics-as-a-service companies. These firms build data products on top of disparate solutions such as Tableau, Snowflake and Redshift, and offer insights to decision-makers in diverse verticals, from finance and marketing to energy and transportation.

Eliminate the pitfalls on your path to public cloud

As organizations look to get smarter and more agile in how they gain value and insight from their data, they are now able to take advantage of a fundamental shift in architecture. In the last decade, as an industry, we have gone from monolithic machines with direct-attached storage to VMs to cloud. The main attraction of cloud is due to its separation of compute and storage – a major architectural shift in the infrastructure layer that changes the way data can be stored and processed.

How to run queries periodically in Apache Hive

In the lifecycle of a data warehouse in production, there are a variety of tasks that need to be executed on a recurring basis. To name a few concrete examples, scheduled tasks can be related to data ingestion (inserting data from a stream into a transactional table every 10 minutes), query performance (refreshing a materialized view used for BI reporting every hour), or warehouse maintenance (executing replication from one cluster to another on a daily basis).