Like many people I was attending the All Hands Meeting in Cardiff last week and the first EGI Technical Forum meeting which was taking place in Amsterdam, home to the EGI headquarters.
However before I had to leave AHM I did attend some interesting sessions on "Sharing, Collaboration and Interfaces for e-Research" featuring "BlogMyData" which Jason has already mentioned.
I have to thank Andrew Richards, Director of the NGS, for giving my presentation in the "Enhancing Community Intelligence for e-Science" workshop which I unfortunately couldn't attend due to being on a plane to Amsterdam! The organiser Alex Voss had asked me to report on some of the statistics that the NGS collects in its usual day-to-day running. This includes the data from the user application forms, the NGS member sites and much more. The invitation to present on these stats was very timely as we have recently released a new service to users on the NGS homepage.
We have made a selection of statisitics publicly available on the NGS website including statistics by research area, institution, NGS usage over time, funding sources and information sources and more. The link can be found on the right hand side of the home page underneath the latest poll.
Meanwhile at the EGI conference in Amsterdam I met up with the EGI dissemination team for the first time as well other dissemination people from several other NGI's. It's amazing no matter how far apart the countries, we all have the same problems and challenges in getting the word out there about grid computing and hunting down those user success stories! Watch out for more user case studies from all over Europe!
Monday, 20 September 2010
Wednesday, 15 September 2010
Bigger data and Faster Horses
Alex Szalay of John Hopkins University and the Sloan Digital Sky Survey spoke at All Hands this morning.
His talk was about the science that can be done when you simply have too much data to store or process and your first task is working out which bits you need to throw away.
Among the many interesting points he made was that, by Amdahl's Law, modern computers are unbalanced if they are used for data-driven research.
CPUs are fast. Modern multi-core CPUs can crunch numbers at extra-ordinary rates. But we gain very little from this if we can't feed them the numbers as fast as they can crunch them.
At best, the numbers have to be loaded from memory and, on the timescales at which a computer works, memory access is slow. At worst, they come from disk and disk access is much, much slower.
Modern CPUs hide the generally sluggishness of memory by keeping data that has been used or may soon be used within a small but very fast caches. As a block of data is transferred from the main memory to cache, nearby blocks are copied too.
Disks and operating systems use a similar approach - whenever a user requests a block of data to be transferred from disk to memory, the blocks that follow it are transferred too.
You only get the advantage of the memory caches and disk access if you are reading the data in one big long stream.
Prof. Szalay likened this to watching the results of a laboratory experiment as it runs. He described computer systems - which he called Data Scopes - designed so that the speed at which data can be accessed is as near as possible to the speed at which it can be processed. You carefully layout your data in the 'scope and just let its comparatively low powered CPUs crunch away.
It is very different from the current approach to High Performance Computing. He ended his talk with a quote attributed to Henry Ford - an example of why 'more of the same' is not always an option:
His talk was about the science that can be done when you simply have too much data to store or process and your first task is working out which bits you need to throw away.
Among the many interesting points he made was that, by Amdahl's Law, modern computers are unbalanced if they are used for data-driven research.
CPUs are fast. Modern multi-core CPUs can crunch numbers at extra-ordinary rates. But we gain very little from this if we can't feed them the numbers as fast as they can crunch them.
At best, the numbers have to be loaded from memory and, on the timescales at which a computer works, memory access is slow. At worst, they come from disk and disk access is much, much slower.
Modern CPUs hide the generally sluggishness of memory by keeping data that has been used or may soon be used within a small but very fast caches. As a block of data is transferred from the main memory to cache, nearby blocks are copied too.
Disks and operating systems use a similar approach - whenever a user requests a block of data to be transferred from disk to memory, the blocks that follow it are transferred too.
You only get the advantage of the memory caches and disk access if you are reading the data in one big long stream.
Prof. Szalay likened this to watching the results of a laboratory experiment as it runs. He described computer systems - which he called Data Scopes - designed so that the speed at which data can be accessed is as near as possible to the speed at which it can be processed. You carefully layout your data in the 'scope and just let its comparatively low powered CPUs crunch away.
It is very different from the current approach to High Performance Computing. He ended his talk with a quote attributed to Henry Ford - an example of why 'more of the same' is not always an option:
If I had asked people what they wanted, they would have said faster horses.
Tuesday, 14 September 2010
It's good to be back at All Hands and catch up with all the Old Hands, and shake New Hands as well.
As usual it is necessary to clone oneself for the parallel sessions, keep track of all the discussions, keep the todo list up to date, sniff around for new things, while simultaneously keeping up with email back home, other projects ticking along at home, documents, proposals, reports, arrangements. The mental equivalent of an octopus. All in a day's work. Hey ho.
As usual it is necessary to clone oneself for the parallel sessions, keep track of all the discussions, keep the todo list up to date, sniff around for new things, while simultaneously keeping up with email back home, other projects ticking along at home, documents, proposals, reports, arrangements. The mental equivalent of an octopus. All in a day's work. Hey ho.
Seen at All Hands
The All Hands meeting exists so that those involved in one branch of e-Intrastructure / e-Science / e-Research / e-Social-Science / e-Whatever can see what those in the other e-*s are up to.
One project that particularly caught my eye was BlogMyData.
Much academic research takes place in corridors, pubs and even - occasionally - in the loo. Researchers will discuss their latest discoveries with colleagues when they bump into one another on the way to somewhere else - and get a new perspective or a new idea as they do so. Call it serendipity at work - or possibly serendipity in the bar of the Dog and Duck.
This is the kind of material that occasionally appears as `Bloggs, Fred (Personal Communcation)' in papers.
BlogMyData extends this chatter about the work to researchers who are in different institutions and so - unless they happen to be at All Hands - are very unlikely to be in the same coffee room, or pub, at the same time.
It allows researchers to post visualisations of the data they are working on blogs which can be read - and commented on - by collaborators. It combines two projects: the Godiva 2 visualisation package from Reading and the LabBlog blogging tool from platform.
This has the big advantage that the data and the conversation will be recorded for future reference unlike, say, a chat in the bar or an unexpected encounter in the gents...
One project that particularly caught my eye was BlogMyData.
Much academic research takes place in corridors, pubs and even - occasionally - in the loo. Researchers will discuss their latest discoveries with colleagues when they bump into one another on the way to somewhere else - and get a new perspective or a new idea as they do so. Call it serendipity at work - or possibly serendipity in the bar of the Dog and Duck.
This is the kind of material that occasionally appears as `Bloggs, Fred (Personal Communcation)' in papers.
BlogMyData extends this chatter about the work to researchers who are in different institutions and so - unless they happen to be at All Hands - are very unlikely to be in the same coffee room, or pub, at the same time.
It allows researchers to post visualisations of the data they are working on blogs which can be read - and commented on - by collaborators. It combines two projects: the Godiva 2 visualisation package from Reading and the LabBlog blogging tool from platform.
This has the big advantage that the data and the conversation will be recorded for future reference unlike, say, a chat in the bar or an unexpected encounter in the gents...
Hello from AHM
A selection of NGS staff are all now settled into conference mode at the UK e-Science All Hands Meeting in a wet Cardiff.
The stand was put up last night and our first demo was held this morning. A big thanks to Jonathan Churchill whose demo of the UI/WMS managed to pull in a substantial crowd despite a quiet start to the event as people continue arrive during the morning.
Jonathan will be doing a follow up demo this lunchtime on the "gLite WMS Enabled NGS Applications Portal" so pop by our stand with your lunch if you are in Cardiff!
The conference proper kicks off this afternoon with the first themes and workshops. I'll be going to the "Sharing, collaboration and interfaces for e-Research" which looks as though it will have some interesting presentations about user tools.
The stand was put up last night and our first demo was held this morning. A big thanks to Jonathan Churchill whose demo of the UI/WMS managed to pull in a substantial crowd despite a quiet start to the event as people continue arrive during the morning.
Jonathan will be doing a follow up demo this lunchtime on the "gLite WMS Enabled NGS Applications Portal" so pop by our stand with your lunch if you are in Cardiff!
The conference proper kicks off this afternoon with the first themes and workshops. I'll be going to the "Sharing, collaboration and interfaces for e-Research" which looks as though it will have some interesting presentations about user tools.
Saturday, 11 September 2010
The Featherstone-Kite Openwork Basketweave Mark Two Gentleman’s Flying Machine
In a shopping centre in the middle of Leeds, not far from the Universities, there is - or was - a big glass case.
In the case is a contraption built from wicker-work, string and cogs and lights and bits of old gramophone. Every so often, it springs to life and whirls around and plays a tune.
It is not an Yorkshire-based competitor for the iPod but a sculpture by Rowland Emmet: "The Featherstone-Kite Openwork Basketweave Mark Two Gentleman’s Flying Machine". Leeds shoppers passing by look at it and think...
We are investigating how to move the important features of the existing INCA monitoring service to a new monitoring service based on WLCG Nagios.
But - as has been said many times - grids are complicated. Which means that the software needed to monitor grids is complicated. Which means that when you start to look at the software, you spend a lot of time staring at a screen and thinking...
So after a week of staring and thinking, here is what we think the bits and pieces of the service are meant to do:
At the core sits Nagios: an open-source monitoring system familiar to many system administrators. It consists of a set of programs called 'plugins' and a scheduler that arranges for these plugins to be run.
A plugin tests if a particular service on a given host is working as expected. Plugins typically return a short message and a status code that means one of: 'OK', 'WARNING', 'CRITICAL' or - if the plugin broke - 'UNKNOWN'. They can also track performance data such as disk usage.
Nagios comes with a set of basic plugins. WLCG Nagios adds a whole raft of Grid specific ones.
In this documentation, plugins within WLCG are referred to as probes.
Next up, a 'configuration generator' called NCG takes data published about a site or set of sites and generates a configuration for Nagios that monitors them.
Statistics and performance metrics generated by the plugins/probes are collected and are delivered via a message bus to a service that stuffs them into a database. A tool called MyEGEE is used to visualise the contents of this database.
If you want to know more...
Staff from STFC and Oxford gave an NGS surgery on WLCG Nagios in late July this year. Their slides describing how WLCG Nagios can be configured and how it has been deployed can be found on the NGS web site.
There is more technical information on twiki.cern.ch in the GridMonitoringNcgOverview and GridMonitoringNcgYaim pages. More information about the plugins/probes can be found on SAMProbesMetrics.
If you want to know more about the NGS R+D activity, we will on on hand at All Hands next week.
In the case is a contraption built from wicker-work, string and cogs and lights and bits of old gramophone. Every so often, it springs to life and whirls around and plays a tune.
It is not an Yorkshire-based competitor for the iPod but a sculpture by Rowland Emmet: "The Featherstone-Kite Openwork Basketweave Mark Two Gentleman’s Flying Machine". Leeds shoppers passing by look at it and think...
What on earth is THAT meant to do?Which is a rather tortuous way of introducing the latest bit of R+D work.
We are investigating how to move the important features of the existing INCA monitoring service to a new monitoring service based on WLCG Nagios.

But - as has been said many times - grids are complicated. Which means that the software needed to monitor grids is complicated. Which means that when you start to look at the software, you spend a lot of time staring at a screen and thinking...
What on earth is THAT meant to do?
So after a week of staring and thinking, here is what we think the bits and pieces of the service are meant to do:
At the core sits Nagios: an open-source monitoring system familiar to many system administrators. It consists of a set of programs called 'plugins' and a scheduler that arranges for these plugins to be run.
A plugin tests if a particular service on a given host is working as expected. Plugins typically return a short message and a status code that means one of: 'OK', 'WARNING', 'CRITICAL' or - if the plugin broke - 'UNKNOWN'. They can also track performance data such as disk usage.
Nagios comes with a set of basic plugins. WLCG Nagios adds a whole raft of Grid specific ones.
In this documentation, plugins within WLCG are referred to as probes.
Next up, a 'configuration generator' called NCG takes data published about a site or set of sites and generates a configuration for Nagios that monitors them.
Statistics and performance metrics generated by the plugins/probes are collected and are delivered via a message bus to a service that stuffs them into a database. A tool called MyEGEE is used to visualise the contents of this database.
If you want to know more...
Staff from STFC and Oxford gave an NGS surgery on WLCG Nagios in late July this year. Their slides describing how WLCG Nagios can be configured and how it has been deployed can be found on the NGS web site.
There is more technical information on twiki.cern.ch in the GridMonitoringNcgOverview and GridMonitoringNcgYaim pages. More information about the plugins/probes can be found on SAMProbesMetrics.
If you want to know more about the NGS R+D activity, we will on on hand at All Hands next week.
Tuesday, 7 September 2010
Social Sciences and Humanities and the NGS Innovation Forum
We are pleased to say that Martin Wynne from the CLARIN project will be presenting at the forthcoming NGS Innovation Forum. We have a number of social scientists who use the NGS and this is an excellent opportunity to demonstrate to the community how the NGS can be of relevance in this research area.
Martins presentation will look at what is necessary at the national level to support CLARIN, a European infrastructure providing services relating to language resources and tools to researchers in the Humanities and Social Sciences.
Continuing with the European theme we are also pleased to announce that Steven Newhouse, Director of EGI, will also be presenting at the event. Steven will be reflecting on the experiences of the first 6 months and provide an overivew for the plans that have now been established for future years.
Registration for the Innovation Forum is now open with full details available on the event page on the NGS website.
Martins presentation will look at what is necessary at the national level to support CLARIN, a European infrastructure providing services relating to language resources and tools to researchers in the Humanities and Social Sciences.
Continuing with the European theme we are also pleased to announce that Steven Newhouse, Director of EGI, will also be presenting at the event. Steven will be reflecting on the experiences of the first 6 months and provide an overivew for the plans that have now been established for future years.
Registration for the Innovation Forum is now open with full details available on the event page on the NGS website.
Subscribe to:
Posts (Atom)