Sat.May 29, 2021 - Fri.Jun 04, 2021

article thumbnail

Turning the page

Cloudera

Today marks the beginning of an exciting new chapter for Cloudera. Cloudera will become a private company with the flexibility and resources to accelerate product innovation, cloud transformation and customer growth. Cloudera will benefit from the operating capabilities, capital support and expertise of Clayton, Dubilier & Rice (CD&R) and KKR – two of the most experienced and successful global investment firms in the world recognized for supporting the growth strategies of the businesses

Cloud 144
article thumbnail

Streaming ETL and Analytics on Confluent with Maritime AIS Data

Confluent

One of the canonical examples of streaming data is tracking location data over time. Whether it’s ride-sharing vehicles, the position of trains on the rail network, or tracking airplanes waking […].

Data 117
Insiders

Sign Up for our Newsletter

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

article thumbnail

Data Engineers of Netflix?—?Interview with Dhevi Rajendran

Netflix Tech

Data Engineers of Netflix?—?Interview with Dhevi Rajendran Dhevi Rajendran This post is part of our “Data Engineers of Netflix” interview series, where our very own data engineers talk about their journeys to Data Engineering @ Netflix. Dhevi Rajendran is a Data Engineer on the Growth Data Science and Engineering team. Dhevi joined Netflix in July 2020 and is one of many Data Engineers who have onboarded remotely during the pandemic.

article thumbnail

Bundle and Distribute Next.js Sites via NPM

Grouparoo

Grouparoo uses Next.js to build our web frontend(s), and we distribute these frontend User Interfaces (UIs) via NPM as packages, e.g. @grouparoo/ui-community. This allows Grouparoo users to choose which UI they want to use (or none) by changing their package.json : Example package.json for a Grouparoo project: { "author" : "Your Name <email@example.com>" , "name" : "grouparoo-application" , "description" : "A Grouparoo Deployment" , "ve

Project 52
article thumbnail

Beyond the Basics of A/B Tests: Highly Innovative Experimentation Tactics You Need to Know

Speaker: Timothy Chan, PhD., Head of Data Science

Are you ready to move beyond the basics and take a deep dive into the cutting-edge techniques that are reshaping the landscape of experimentation? 🌐 From Sequential Testing to Multi-Armed Bandits, Switchback Experiments to Stratified Sampling, Timothy Chan, Data Science Lead, is here to unravel the mysteries of these powerful methodologies that are revolutionizing how we approach testing.

article thumbnail

Validations – Cloudera Support’s Predictive Alerting Program

Cloudera

Cloudera Support’s cluster validations proactively identify known problem signatures contained in customers’ diagnostic data with the goal of increasing cluster health, performance, and overall stability. Cluster validations are included in a customer’s enterprise subscription at no additional cost. All customers with access to the Support case portal will also be able to take advantage of cluster validations.

article thumbnail

Are We There Yet? The Query Your Database Can’t Answer

Confluent

What if I told you there is a query your database can’t answer? That would probably surprise you. With decades of effort behind them, databases are one of the most […].

More Trending

article thumbnail

Monte Carlo Brings Data Observability to Data Lakes with New Databricks Integration

Monte Carlo

As companies leverage more and more data to drive decision-making and maintain their competitive edge, it’s crucial that this data is accurate and reliable. With the new Databricks integration from Monte Carlo, teams working in data lakes can finally trust their data through end-to-end data observability and automated lineage of their entire data ecosystem.

article thumbnail

Modernizing Data Pipelines using Cloudera Data Platform – Part 1

Cloudera

Data pipelines are in high demand in today’s data-driven organizations. As critical elements in supplying trusted, curated, and usable data for end-to-end analytic and machine learning workflows, the role of data pipelines is becoming indispensable. To keep up, data pipelines are being vigorously reshaped with modern tools and techniques. At Cloudera, we recently introduced several cutting-edge innovations in our Cloudera Data Engineering experience (CDE) as part of our Enterprise Data Cloud pro

article thumbnail

Detecting Patterns of Behaviour in Streaming Maritime AIS Data with Confluent

Confluent

This is part two in a blog series on streaming a feed of AIS maritime data into Apache Kafka® using Confluent and how to apply it for a variety of […].

Kafka 57
article thumbnail

Billions of Personal Interactions

Teradata

To meet evolving customer demands banks must find ways to manage billions of hyper-personalized interactions at low cost. This demands the right tech & operating model.

Banking 52
article thumbnail

From Developer Experience to Product Experience: How a Shared Focus Fuels Product Success

Speaker: Anne Steiner and David Laribee

As a concept, Developer Experience (DX) has gained significant attention in the tech industry. It emphasizes engineers’ efficiency and satisfaction during the product development process. As product managers, we need to understand how a good DX can contribute not only to the well-being of our development teams but also to the broader objectives of product success and customer satisfaction.

article thumbnail

The Techniques and Technologies Bringing Agility to Enterprise Data

DataKitchen

The post The Techniques and Technologies Bringing Agility to Enterprise Data first appeared on DataKitchen.

article thumbnail

Top 30 IoT-based Projects for Beginners in 2023

ProjectPro

Are you a final-year student or a beginner looking to improve your skills in IoT technology? If yes, this blog has some of the best and most exciting IoT projects for you! You will find various IoT-based projects relevant to different industries that can help you land your dream job in data science! IoT is likely to grow from 8.74 billion in 2020 to more than 13 billion in 2023, according to Statista Research Department.

Project 40
article thumbnail

Rockset Converged Index Adds Clustered Search Index for 70% Query Latency Reduction

Rockset

In this blog, we will describe a new storage format that we adopted for our search index, one of the indexes in Rockset’s Converged Index. This new format reduced latencies for common queries by as much as 70% and the size of the search index by about 20%. As described in our Converged Index blog, we store every column of every document in a row-based store, column-based store, and a search index.

article thumbnail

It Just Got a Lot Easier to Offload Data From Vantage to Cloud Storage

Teradata

Learn about the capabilities and benefits of NOS WRITE -- the latest offering within the Native Object Store feature, which was released in early 2020.

article thumbnail

Peak Performance: Continuous Testing & Evaluation of LLM-Based Applications

Speaker: Aarushi Kansal, AI Leader & Author and Tony Karrer, Founder & CTO at Aggregage

Software leaders who are building applications based on Large Language Models (LLMs) often find it a challenge to achieve reliability. It’s no surprise given the non-deterministic nature of LLMs. To effectively create reliable LLM-based (often with RAG) applications, extensive testing and evaluation processes are crucial. This often ends up involving meticulous adjustments to prompts.

article thumbnail

Making Data Pipelines Self-Serve For Everyone With Shipyard

Data Engineering Podcast

Summary Every part of the business relies on data, yet only a small team has the context and expertise to build and maintain workflows and data pipelines to transform, clean, and integrate it. In order for the true value of your data to be realized without burning out your engineers you need a way for everyone to get access to the information they care about.

article thumbnail

Apache Ozone Metadata Explained

Cloudera

Apache Ozone is a distributed object store built on top of Hadoop Distributed Data Store service. It can manage billions of small and large files that are difficult to handle by other distributed file systems. As an important part of achieving better scalability, Ozone separates the metadata management among different services: . Ozone Manager (OM) service manages the metadata of the namespace such as volume, bucket and keys.

article thumbnail

Getting Started with Real-Time Analytics on MySQL Using Rockset

Rockset

MySQL and PostgreSQL are widely used as transactional databases. When it comes to supporting high-scale and analytical use cases, you may often have to tune and configure these databases, which leads to a higher operational burden. Some challenges when doing analytics on MySQL and Postgres include: running a large number of concurrent queries/users working with large data sizes needing to define and manage tons of indexes.

MySQL 40
article thumbnail

Exploratory Data Analysis in Python-Stop, Drop and Explore

ProjectPro

"Exploratory data analysis is an attitude, a state of flexibility, a willingness to look for those things that we believe are not there, as well as the things we believe might be there. “ - quoted in Exploratory Data Analysis Tukey PDF on Nonparametric Statistical Data Modeling. This data science blog will discover what is exploratory data analysis (EDA), the importance of performing EDA when solving data science problems, the various exploratory data analysis techniques that one can use w

article thumbnail

Entity Resolution Checklist: What to Consider When Evaluating Options

Are you trying to decide which entity resolution capabilities you need? It can be confusing to determine which features are most important for your project. And sometimes key features are overlooked. Get the Entity Resolution Evaluation Checklist to make sure you’ve thought of everything to make your project a success! The list was created by Senzing’s team of leading entity resolution experts, based on their real-world experience.

article thumbnail

Build Your Analytics With A Collaborative And Expressive SQL IDE Using Querybook

Data Engineering Podcast

Summary SQL is the most widely used language for working with data, and yet the tools available for writing and collaborating on it are still clunky and inefficient. Frustrated with the lack of a modern IDE and collaborative workflow for managing the SQL queries and analysis of their big data environments, the team at Pinterest created Querybook. In this episode Justin Mejorada-Pier and Charlie Gu share the story of how the initial prototype for a data catalog ended up as one of their most widel

SQL 100
article thumbnail

How Airbnb Standardized Metric Computation at Scale

Airbnb Tech

Metric Infrastructure with Minerva @ Airbnb Part II: The six design principles of Minerva compute infrastructure By: Amit Pahwa , Cristian Figueroa , Donghan Zhang , Haim Grosman , John Bodley , Jonathan Parks , Maggie Zhu , Philip Weiss , Robert Chang , Shao Xie , Sylvia Tomiyama , Xiaohui Sun Introduction As described in the first post of this series, Airbnb invested significantly in building Minerva, a single source of truth metric platform that standardizes the way business metrics are creat

article thumbnail

Announcing the 2021 Data Platform Trends Report

Monte Carlo

In this report, we’ll dive into the latest data platform trends and challenges, as well as outline the key factors teams should consider when building their solutions, including: Data discovery Data mesh Data democratization Data observability And many others. We’ll also highlight how teams at Uber, Nike, Pinterest, and other companies leading with data apply these concepts at scale to drive value and trustworthy insights for the business.

Data 52
article thumbnail

20 Solved End-to-End Big Data Projects with Source Code

ProjectPro

Ace your big data interview by adding some unique and exciting Big Data projects to your portfolio. This blog lists over 20 big data projects you can work on to showcase your big data skills and gain hands-on experience in big data tools and technologies. You will find several big data projects depending on your level of expertise- big data projects for students, big data projects for beginners, etc.

article thumbnail

How to Build an Experimentation Culture for Data-Driven Product Development

Speaker: Margaret-Ann Seger, Head of Product, Statsig

Experimentation is often seen as an aspirational practice, especially at smaller, fast-moving companies who are strapped for time and resources. So, how can you get your team making decisions in a more data-driven way while continuing to remain lean and maintaining ship velocity? In this webinar, Margaret-Ann Seger, Head of Product at Statsig, will teach you how to build an experimentation culture from the ground-up, graduating from just getting started with data-driven development to operating