Top Data Engineering Digest Big Data Skills Big Data Content for 2018

2018

Functional Data Engineering — a modern paradigm for batch data processing

Maxime Beauchemin

JANUARY 7, 2018

Batch data processing — historically known as ETL — is extremely challenging. It’s time-consuming, brittle, and often unrewarding. Not only that, it’s hard to operate, evolve, and troubleshoot. In this post, we’ll explore how applying the functional programming paradigm to data engineering can bring a lot of clarity to the process. This post distills fragments of wisdom accumulated while working at Yahoo, Facebook, Airbnb and Lyft, with the perspective of well over a decade of data warehousing

Data Engineering

Data Engineering Data Engineer Data Process Process

Open-Source Data Warehousing – Druid, Apache Airflow & Superset

Simon Späti

NOVEMBER 28, 2018

These days, everyone talks about open-source. However, this is still not common in the Data Warehouse (DWH) field. Why is this? In my recent blog, I researched OLAP technologies, for this post I chose some open-source technologies and used them together to build a full data architecture for a Data Warehouse system. I went with Apache Druid for data storage, Apache Superset for querying and Apache Airflow as a task orchestrator.

Data Warehouse

Data Warehouse Data Storage Data Architecture Architecture

Join 16,000+

Insiders

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Webinars

Demystifying DAPs: A Practical Guide to Digital Adoption Success

The AI Superhero Approach to Product Management

MORE WEBINARS

Trending Sources

Octopai: Metadata Management for Better Business Intelligence with Amnon Drori - Episode 28

Data Engineering Podcast

APRIL 22, 2018

Summary The information about how data is acquired and processed is often as important as the data itself. For this reason metadata management systems are built to track the journey of your business data to aid in analysis, presentation, and compliance. These systems are frequently cumbersome and difficult to maintain, so Octopai was founded to alleviate that burden.

Business Intelligence

Business Intelligence Metadata Management Data Governance

Webinars

Demystifying DAPs: A Practical Guide to Digital Adoption Success

The AI Superhero Approach to Product Management

MORE WEBINARS

Observability at Scale: Building Uber’s Alerting Ecosystem

Uber Engineering

NOVEMBER 20, 2018

Uber’s software architectures consists of thousands of microservices that empower teams to iterate quickly and support our company’s global growth. These microservices support a variety of solutions, such as mobile applications, internal and infrastructure services, and products along with complex … The post Observability at Scale: Building Uber’s Alerting Ecosystem appeared first on Uber Engineering Blog.

Building

Building Architecture Engineering

The AI Superhero Approach to Product Management

Speaker: Conrado Morlan

In this engaging and witty talk, we’ll explore how artificial intelligence can transform the daily tasks of product managers into streamlined, efficient processes. Using the lens of a superhero narrative, we’ll uncover how AI can be the ultimate sidekick, aiding in decision-making, enhancing productivity, and boosting innovation. Attendees will leave with practical tools and actionable insights, motivated to embrace AI and leverage its potential in their work. 🦸 🏢 Key objectives:

Management

Netflix OSS and Spring Boot?—?Coming Full Circle

Netflix Tech

DECEMBER 18, 2018

Netflix OSS and Spring Boot?—?Coming Full Circle Taylor Wicksell, Tom Cellucci, Howard Yuan, Asi Bross, Noel Yap, and David Liu In 2007, Netflix started on a long road towards fully operating in the cloud. Much of Netflix’s backend and mid-tier applications are built using Java, and as part of this effort Netflix engineering built several cloud infrastructure libraries and systems?

Java

Java Cloud AWS Government

Do These Things if you Want to Succeed as an HR Professional

U-Next

JANUARY 9, 2018

Success in today’s businesses has taken several meanings. Apart from just pay hikes and promotions, success has gotten new dimensions that have been of very recent origins. Today, success has become synonymous with happiness at a workplace, challenging tasks, compensatory rewards, incentives, authoritative job profiles, influential role, and more. The current talent pools in organizations have become wiser and more mature than their previous generation counterparts.

Recruitment

Recruitment Technology Management IT

Bringing AIOps to Machine Learning & Analytics

Cloudera

AUGUST 31, 2018

Two years ago I founded Hyperpilot with the mission to enable autopilot for container infrastructure. We learned a lot about data center automation based on real-time application and diagnostic feedback using applied machine learning. Last month, I joined Cloudera along with former team members Xiaoyun Zhu and Che-Yuan Liang to bring our expertise in intelligent automation to Cloudera’s modern platform for machine learning and analytics.

Machine Learning

Machine Learning Utilities Cloud Architecture

More Trending

Bringing AIOps to Machine Learning & Analytics

Cloudera

AUGUST 31, 2018

Machine Learning

Machine Learning Utilities Cloud Architecture

Creating Multi-language NLP Pipelines with Apache Spark

Domino Data Lab: Data Engineering

DECEMBER 22, 2018

In this guest post, Holden Karau , Apache Spark Committer , provides insights on how to create multi-language pipelines with Apache Spark and avoid rewriting spaCy into Java. She has already written a complementary blog post on using spaCy to process text data for Domino. Karau is a Developer Advocate at Google as well as a co-author on High Performance Spark and Learning Spark.

Java

Java Coding Process Machine Learning

Live Dashboards on Streaming Data - A Tutorial Using Amazon Kinesis and Rockset

Rockset

DECEMBER 20, 2018

We live in a world where diverse systems—social networks, monitoring, stock exchanges, websites, IoT devices—all continuously generate volumes of data in the form of events, captured in systems like Apache Kafka and Amazon Kinesis. One can perform a wide variety of analyses, like aggregations, filtering, or sampling, on these event streams, either at the record level or over sliding time windows.

AWS

AWS Kafka Data Ingestion Data

One Audio Sequencer to Rule Them All

Pandora Engineering

DECEMBER 5, 2018

Photo credit: Carol Yepes Last month Pandora announced a public podcast beta in conjunction with the Podcast Genome Project. This rollout introduced many exciting features to our current mobile application offerings, including fully integrated and native podcast support. Ironically, one of the most interesting features and perhaps our biggest engineering win with this iteration is something that’s transparent to our end users: the inclusion of a new audio playback sequencer used exclusively for

Media

Media Algorithm Coding Data Science

Open Source: November Review - Maintainer training, new releases and more

Zalando Engineering

DECEMBER 5, 2018

Project Highlights ExternalDNS version 0.5.9 is ready for testing. This project allows you to control DNS records dynamically via Kubernetes resources in a DNS provider-agnostic way. ExternalDNS also successfully made its way to the Kubernetes Incubator. Check out the list of changes in this new release. Zalando-Incubator welcomed two brand new open source projects 1) Darty - a data dependency manager for data science projects.

PostgreSQL

PostgreSQL Java Machine Learning Deep Learning

Provide Real Value in Your Applications with Data and Analytics

The complexity of financial data, the need for real-time insight, and the demand for user-friendly visualizations can seem daunting when it comes to analytics - but there is an easier way. With Logi Symphony, we aim to turn these challenges into opportunities. Our platform empowers you to seamlessly integrate advanced data analytics, generative AI, data visualization, and pixel-perfect reporting into your applications, transforming raw data into actionable insights.

Raw Data

Announcing my session at #SQLBits - Azure Databricks

Advancing Analytics: Data Engineering

DECEMBER 3, 2018

Simon Whiteley and I will be back at #SQLBits 2019 talking about hashtag#DataEngineering and #DataScience in Databricks. We will look at #ApacheSpark #Python #Engineering & #MachineLearning in this full day training day. Register Now Have you looked at Azure DataBricks yet? No! Then you need to. Why you ask, there are many reasons. The number 1, knowing how to use Apache Spark will earn you more money.

Data Science

Data Science Machine Learning Python Data Pipeline

OLAP, what’s coming next?

Simon Späti

NOVEMBER 23, 2018

Are you on the lookout for a replacement for the Microsoft Analysis Cubes, are you looking for a big data OLAP system that scales ad libitum, do you want to have your analytics updated even real-time? In this blog, I want to show you possible solutions that are ready for the future and fits into existing data architecture. What is OLAP? OLAP is an acronym for Online Analytical Processing.

Big Data

Big Data Data Architecture Architecture Systems

Simplifying Continuous Data Processing Using Stream Native Storage In Pravega with Tom Kaitchuck - Episode 63

Data Engineering Podcast

DECEMBER 31, 2018

Summary As more companies and organizations are working to gain a real-time view of their business, they are increasingly turning to stream processing technologies to fullfill that need. However, the storage requirements for continuous, unbounded streams of data are markedly different than that of batch oriented workloads. To address this shortcoming the team at Dell EMC has created the open source Pravega project.

Lambda Architecture

Lambda Architecture Process Data Process Kafka

Maximizing Process Performance with Maze, Uber’s Funnel Visualization Platform

Uber Engineering

AUGUST 16, 2018

At Uber, we spend a considerable amount of resources making the driver sign-up experience as easy as possible. At Uber’s scale, even a one percent increase in the rate of sign-ups to first trips (the driver conversion rate) carries a … The post Maximizing Process Performance with Maze, Uber’s Funnel Visualization Platform appeared first on Uber Engineering Blog.

Process

Process Engineering Architecture Data

Entity Resolution: Your Guide to Deciding Whether to Build It or Buy It

Adding high-quality entity resolution capabilities to enterprise applications, services, data fabrics or data pipelines can be daunting and expensive. Organizations often invest millions of dollars and years of effort to achieve subpar results. This guide will walk you through the requirements and challenges of implementing entity resolution. By the end, you'll understand what to look for, the most common mistakes and pitfalls to avoid, and your options.

Netflix Information Security: Preventing Credential Compromise in AWS

Netflix Tech

NOVEMBER 28, 2018

by Will Bengtson Previously we wrote about a method for detecting credential compromise in your AWS environment. The methodology focused on a continuous learning model and first use principle. This solution still is reactive in nature?—?we only detect credential compromise after it has already happened. Even with detection capabilities, there is a risk that exposed credentials can provide access to sensitive data and/or the ability to cause damage in our environment.

AWS

AWS Metadata Amazon Web Services Cloud

Making slow queries fast with composite indexes in MySQL

nodeSWAT

AUGUST 21, 2018

Making slow queries fast using composite indexes in MySQL This post expects some basic knowledge of SQL. Examples were made using MySQL 5.7.18 and run on my mid 2014 Macbook Pro. Query execution times are based on multiple executions so index caching can kick in. The use-case came from a real application and the solution is used in production. So you have inserted preliminary data to your database and run a simple COUNT(*) query against it with a simple WHERE clause and… the spinner is still run

MySQL

MySQL Datasets Database SQL

Meet the newest Data Superheros: The Sixth Annual Data Impact Awards Finalists Are…

Cloudera

AUGUST 28, 2018

Drum roll… Starting from well over 100 nominations, we are excited to announce the finalists for this year’s Data Impact Awards ! Each year, nominees have raised the bar, and this year is no exception. The level of impact that organizations have shown and the variety of use cases are inspiring. From AI models that power retail customer decision engines to utility meter analysis that disables underperforming gas turbines, these finalists demonstrate how machine learning and analytics have become

Pharmaceutical

Pharmaceutical Telecommunication Consulting Food

Data Science vs Engineering: Tension Points

Domino Data Lab: Data Engineering

DECEMBER 15, 2018

This blog post provides highlights and a full written transcript from the panel, “ Data Science Versus Engineering: Does It Really Have To Be This Way? ” with Amy Heineike , Paco Nathan , and Pete Warden at Domino HQ. Topics discussed include the current state of collaboration around building and deploying models, tension points that potentially arise, as well as practical advice on how to address these tension points.

Data Science

Data Science Engineering Data Building

Generative AI Deep Dive: Advancing from Proof of Concept to Production

Speaker: Maher Hanafi, VP of Engineering at Betterworks & Tony Karrer, CTO at Aggregage

Executive leaders and board members are pushing their teams to adopt Generative AI to gain a competitive edge, save money, and otherwise take advantage of the promise of this new era of artificial intelligence. There's no question that it is challenging to figure out where to focus and how to advance when it’s a new field that is evolving everyday. 💡 This new webinar featuring Maher Hanafi, VP of Engineering at Betterworks, will explore a practical framework to transform Generative AI pr

Data Collection

Recap of Hadoop News for July 2018

ProjectPro

AUGUST 1, 2018

News on Hadoop - July 2018 Hadoop data governance services surface in wake of GDPR.TechTarget.com, July 2, 2018. GDPR has turned out to be a strong motivator that would bring greater governance to big data. At the recent DataWorks Summit 2018 , though most of the attention was focussed on how Hadoop pioneer Hortonworks is all set to expand its service in the cloud, there was great interest and importance put on managing data privacy as well.

Hadoop

Hadoop Pharmaceutical Healthcare Data Lake

Programming Best Practices For Data Science

Dataquest

JUNE 8, 2018

The data science life cycle is generally comprised of the following components: data retrieval data cleaning data exploration and visualization statistical or predictive modeling While these components are helpful for understanding the different phases, they don’t help us think about our programming workflow. Often, the entire data science life cycle ends up as an arbitrary mess of notebook cells in either a Jupyter Notebook or a single messy script.

Data Science

Data Science Programming Data Data Pipeline

#NoEstimates

Zalando Engineering

OCTOBER 31, 2018

Why I advocate a practice of no estimates as a software engineer Before I get to the topic, I would like to clarify one thing: I don’t want to ban estimations generally from software development, as there are good and solid reasons for it. In a nutshell, business needs to be predictable. I want to show a software developer's view on how to reduce or even get rid of endless estimations meetings with doubtful outcomes.

Software Engineer

Software Engineer Software Engineering Coding Building

AI at the Forefront of Digital Transformation Process in 2018

InData Labs

MAY 7, 2018

Digital Transformation Definition Digital transformation has been a big topic for a few years now, and it has many definitions. From a business perspective, digital transformation is about leveraging digital technologies to improve processes, competencies, and business models. It is also about changing the culture of the company because it requires letting go of old.

Process

Process Technology IT Data Science

Beyond the Basics of A/B Tests: Highly Innovative Experimentation Tactics You Need to Know

Speaker: Timothy Chan, PhD., Head of Data Science

Are you ready to move beyond the basics and take a deep dive into the cutting-edge techniques that are reshaping the landscape of experimentation? 🌐 From Sequential Testing to Multi-Armed Bandits, Switchback Experiments to Stratified Sampling, Timothy Chan, Data Science Lead, is here to unravel the mysteries of these powerful methodologies that are revolutionizing how we approach testing.

Data Science

New on Cloud Academy: Machine Learning on Google Cloud and AWS, Big Data Analytics, Terraform, and more

Cloud Academy

MAY 2, 2018

A 2017 IDC White Paper “recommend[s] that organizations that want to get the most out of cloud should train a wide range of stakeholders on cloud fundamentals and provide deep training to key technical teams ” (emphasis ours). Regular readers of the Cloud Academy blog know we’ve been talking about this for a long time. Future-proofing your organization requires technical excellence, collective experience, business context, and shared understanding.

Google Cloud

Google Cloud Machine Learning Big Data AWS

Continuously Query Your Time-Series Data Using PipelineDB with Derek Nelson and Usman Masood - Episode 62

Data Engineering Podcast

DECEMBER 23, 2018

Summary Processing high velocity time-series data in real-time is a complex challenge. The team at PipelineDB has built a continuous query engine that simplifies the task of computing aggregates across incoming streams of events. In this episode Derek Nelson and Usman Masood explain how it is architected, strategies for designing your data flows, how to scale it up and out, and edge cases to be aware of.

PostgreSQL

PostgreSQL Kafka Data Engineering Data Engineer

Databook: Turning Big Data into Knowledge with Metadata at Uber

Uber Engineering

AUGUST 3, 2018

From driver and rider locations and destinations, to restaurant orders and payment transactions, every interaction on Uber’s transportation platform is driven by data. Data powers Uber’s global marketplace, enabling more reliable and seamless user experiences across our products for riders, … The post Databook: Turning Big Data into Knowledge with Metadata at Uber appeared first on Uber Engineering Blog.

Metadata

Metadata Big Data Transportation Data

Cache warming: Agility for a stateful service

Netflix Tech

DECEMBER 4, 2018

by Deva Jayaraman , Shashi Madappa , Sridhar Enugula , and Ioannis Papapanagiotou EVCache has been a fundamental part of the Netflix platform (we call it Tier-1), holding Petabytes of data. Our caching layer serves multiple use cases from signup, personalization, searching, playback, and more. It is comprised of thousands of nodes in production and hundreds of clusters all of which must routinely scale up due to the increasing growth of our members.

AWS

AWS Metadata Architecture Kafka

Demystifying DAPs: A Practical Guide to Digital Adoption Success

Speaker: Pulkit Agrawal

Digital Adoption Platforms (DAPs) are revolutionizing the way organizations interact with and optimize their software applications. As digital transformation continues to accelerate, DAPs have become essential tools for enhancing user engagement and software efficiency. This session is your guide into the robust world of DAPs, exploring their origins, evolution, and the current trends shaping their development.

Certification

Concurrency, MySQL and Node.js: A journey of discovery

nodeSWAT

FEBRUARY 5, 2018

Our story begins like so many others with a code loving protagonist — someone we all can relate to. His days are largely filled with designing code, writing code and reading about code — keeping clients happy while learning and having fun. This has been going on for years now with both MySQL and Node.js among others and as such our protagonist considers himself quite proficient with both those technologies.

MySQL

MySQL Database Programming Coding

Data Engineering is Critical to Big Data Success

Cloudera

JANUARY 12, 2018

I mentioned in an earlier blog titled, “Staffing your big data team, ” that data engineers are critical to a successful data journey. That said, most companies that are early in their journey lack a dedicated engineering group. And the longer it takes to put a team in place, the likelier it is that your big data project will stall. The data engineering team is responsible for collecting and ingesting batch and stream-oriented data, inventorying the data, working through ingest bottlenecks, and d

Big Data

Big Data Data Engineering Data Engineer Engineering

Collaboration Between Data Science and Data Engineering: True or False?

Domino Data Lab: Data Engineering

NOVEMBER 18, 2018

This blog post includes candid insights about addressing tension points that arise when people collaborate on developing and deploying models. Domino’s Head of Content sat down with Don Miner and Marshall Presser to discuss the state of collaboration between data science and data engineering. The blog post provides distilled insights, audio clips, excerpted quotes as well as the full audio and written transcript.

Data Science

Data Science Data Engineering Data Engineer Engineering

Recap of Hadoop News for June 2018

ProjectPro

JULY 3, 2018

News on Hadoop - June 2018 RightShip uses big data to find reliable vessels.HoustonChronicle.com,June 15, 2018. RightShip is using IBM’s predictive big data analytics platform to calculate the likelihood of compliance or mechanical troubles that an individual merchant ship will experience within the next year.It also leverages big data to analyse carbon emissions and vessel efficiency.

Hadoop

Hadoop Big Data Data Mining Government

Deliver Mission Critical Insights in Real Time with Data & Analytics

In the fast-moving manufacturing sector, delivering mission-critical data insights to empower your end users or customers can be a challenge. Traditional BI tools can be cumbersome and difficult to integrate - but it doesn't have to be this way. Logi Symphony offers a powerful and user-friendly solution, allowing you to seamlessly embed self-service analytics, generative AI, data visualization, and pixel-perfect reporting directly into your applications.

Data Analytics

2018

Functional Data Engineering — a modern paradigm for batch data processing

Open-Source Data Warehousing – Druid, Apache Airflow & Superset

Webinars

Trending Sources

Octopai: Metadata Management for Better Business Intelligence with Amnon Drori - Episode 28

Webinars

Observability at Scale: Building Uber’s Alerting Ecosystem

The AI Superhero Approach to Product Management

Netflix OSS and Spring Boot?—?Coming Full Circle

Do These Things if you Want to Succeed as an HR Professional

Bringing AIOps to Machine Learning & Analytics

Sign up to get articles personalized to your interests!

More Trending

Bringing AIOps to Machine Learning & Analytics

Creating Multi-language NLP Pipelines with Apache Spark

Live Dashboards on Streaming Data - A Tutorial Using Amazon Kinesis and Rockset

One Audio Sequencer to Rule Them All

Open Source: November Review - Maintainer training, new releases and more

Provide Real Value in Your Applications with Data and Analytics

Announcing my session at #SQLBits - Azure Databricks

OLAP, what’s coming next?

Simplifying Continuous Data Processing Using Stream Native Storage In Pravega with Tom Kaitchuck - Episode 63

Maximizing Process Performance with Maze, Uber’s Funnel Visualization Platform

Entity Resolution: Your Guide to Deciding Whether to Build It or Buy It

Netflix Information Security: Preventing Credential Compromise in AWS

Making slow queries fast with composite indexes in MySQL

Meet the newest Data Superheros: The Sixth Annual Data Impact Awards Finalists Are…

Data Science vs Engineering: Tension Points

Generative AI Deep Dive: Advancing from Proof of Concept to Production

Recap of Hadoop News for July 2018

Programming Best Practices For Data Science

#NoEstimates

AI at the Forefront of Digital Transformation Process in 2018

Beyond the Basics of A/B Tests: Highly Innovative Experimentation Tactics You Need to Know

New on Cloud Academy: Machine Learning on Google Cloud and AWS, Big Data Analytics, Terraform, and more

Continuously Query Your Time-Series Data Using PipelineDB with Derek Nelson and Usman Masood - Episode 62

Databook: Turning Big Data into Knowledge with Metadata at Uber

Cache warming: Agility for a stateful service

Demystifying DAPs: A Practical Guide to Digital Adoption Success

Concurrency, MySQL and Node.js: A journey of discovery

Data Engineering is Critical to Big Data Success

Collaboration Between Data Science and Data Engineering: True or False?

Recap of Hadoop News for June 2018

Deliver Mission Critical Insights in Real Time with Data & Analytics

Stay Connected