Extreme Data

Longform
Image: Stcokfresh

15 May 2014

High performance centred
The Irish Centre for High-End Computing (ICHEC), founded in 2005, is Ireland’s national centre for High-Performance Computing (HPC) with resources, support, education and training for researchers in third-level institutions and technology transfer to industry. It is also the pivot for Ireland’s engagement in HPC and related fields throughout the EU and globally. Dr Alistair McKinstry heads up its environmental sciences acidity, which embraces climatology and remote satellite observation to name two topical fields.

Happy to use the V terms, McKinstry says that volume and veracity are the challenging ones in most of his work, although velocity has begun to be an issue with satellite observations. “Our dominant work for the past few years has been in contribution to the UN Intergovernmental Panel on Climate Change (IPCC) reports. We have been working since 2008 with Met Eireann and 14 other bodies around Europe on a new climate model called EC Earth. That in turn contributes to the Fifth Coupled Model Intercomparison Project (CMIP5), which is a coordinated set of experiments on the global climate models built by labs around the world. The various models do not give the same answers, not surprisingly, so the project looks at how and where and why they differ and how close to reality are they or have they been.”

Essentially, a total of 25 agencies around the world each run the same carefully defined experimental scenarios on their specific climate models every few years and then compare and analyse the results. In terms of data volume, the last full CMIP count showed 3.3 petabytes in the total dataset with 100 participating. A knowledgeable colleague recently forecast that by 2020 each of those experiments will be working with an Exabyte of data or even more, McKinstry says.

“It has become clear to everyone involved in this and similar projects globally that we need the skills of what we now call a ‘data scientist’ to develop our computer models further. Take a small example. When the reasons for divergences between results from different climate models were investigated it was discovered that ‘sea temperature’ might mean different things in different models, or soil temperature might be sampled at different depths. So you can have massive volumes of accurate data yet models and analysis can produce different results because of relatively small but unrecognised factors. They could be in the scientific ‘facts’ or in the analytical processes. In climatology, as in many other spheres of science, we have made enormous progress. But we have also come to recognise at least some of the areas in which we must improve.”

Better decision making
“The whole point for collecting data is so that you can derive some value from it, and in the case of big data usually that is to help us make better decisions,” says Dr Randy Cogill of IBM Research Ireland. Assistant Professor in the department of Systems and Information Engineering at University of Virginia, the main thrust of his research group is the development of optimisation and control techniques for improving system performance.

“Today that presents us with a whole new set of difficult challenges, notably differences in what has conventionally been done in statistics and decision sciences. This is data science. As we work with larger and larger data sets, with much more variety in them, we see that there is lot of garbage in them as well as real information,” Cogill says.

It has become clear to everyone involved in this and similar projects globally that we need the skills of what we now call a ‘data scientist’ to develop our computer models further, Dr Alistair McKinstry, Irish Centre for High-End Computing

“We are constantly developing hardware and software to manage larger and larger data sets, but making sense of them is the challenge — particularly when they are corrupted with more and more garbage! I see two classes of challenge — one is to drive better decision making and the other is to derive real insights rather than spurious correlations. That is where data science is of growing importance, in managing the analysis and indeed in keeping an eye firmly also on the value of the insights. Because there will always be lots of things in the data that are true, but actually useless.”

The IBM centre in Dublin is focussed on its Smarter Cities project, which Dr Cogill believes is a potentially huge area for ever-bigger data and analytics. “We can take data from multiple sensors — in fact every single person or car could be a multi-sensor platform. All of these, together with fixed point and passive sensors, can be a dynamic data infrastructure. It is in the fusion of all of those possible data sets that we can gain insights and value. One huge part of the challenge is that much of this data is always going to be noisy. Then you have to ask, maybe what we are looking at is just noise?”

The extreme data that will ultimately be produced by Smarter Cities projects is a great example, Dr Cogill says, because there is so much opportunity to make huge gains. “Urban infrastructure, for example, has simply never been highly optimised. But if you think of the city as an ecosystem then a constant flow of open, shared data would be the basis for constant improvement and development for citizens, systems and urban management.”

 

Read More:


Back to Top ↑