As you might know, many of us are meeting next week in Baltimore (physically and virtually) to discuss portal development in Year 2 (which we plan to dive into right after the second beta release of the Data Discovery Tool at the end of September). This meeting is part of an effort across all our science initiatives to ensure that our development is science-driven. Consequently, the big focus of this meeting will be science use cases. Gretchen Greene has assembled a great set of use cases that were developed by a number of people reaching back to the NVO days and spanning to more recent examples from the VAO and the IVOA. This week I spent some time with these use cases to try to pull out some common features and needs. It was an inspiring and worthwhile exercise--I highly recommend it.
Doug Tody made the comment this week that considering these use cases as they are laid out on that page makes this meeting more of one about the VAO architecture. I believe he was making reference to the fact that our first year's effort for the portal was about data discovery, resulting in a nicely focused tool that includes some new techniques for quickly finding and selecting datasets. The use cases, on the other hand, go beyond discovery. So, yes, I have to agree with him: addressing these use cases does mean addressing the VAO architecture.He also said in the same breath, "I'm not sure what the Portal is." As we take a step back to consider what we've accomplished in the first year and what we'll aim for in the second, it's a good time to try re-answering that question. We should do it in terms of science use cases--what it is we want astronomers to accomplish--and, yes, this means addressing some of the essentials of the VAO architecture.
I want to suggest that we see the Portal as a web-based entryway to a collection of tools that work together. What you can do in the portal goes beyond data discovery; that is, the Data Discovery Tool is just one of those tools. The Cross Comparison Tool should be another (actually quite a critical one, as I hope to discuss in a later post). The TAP client should have a presence there as well. And what can you accomplish in the portal? Those use cases, of course.
Updated 9/6: Given the discussion below, I should note that I meant to say that the Portal--in my definition--is explicitly browser-based (with programmatic web service interfaces possible, too).
In this vision, the Portal is not a monolithic tool that does everything; each tool specializes in some task, and over time we can add more tools. It's worth noting though that like all good use cases, ours say nothing about "tools." Clearly, to solve those problems, one will have to use several tools together. Thus, a key aim of the Portal is allowing the tools to work together: the user should be able to easily discover and create lists of datasets and sources with the Discovery Tool and join them together with the Cross-Comparison Tool. It would be great if that collaboration of tools were so quick and smooth, that the user may not realize that two different tools were used. Of course, this kind of collaboration extends to the desktop: the user should be able to send selected catalog records to Topcat and a list of SEDs to Iris.
A key question, then, we need to answer in the design of the Portal is what is needed to allow these tools to work together? This, some of you will recall, was a key problem that had to be tackled when we were developing the NVO portal. Fortunately, we have a few more technological options available to us today than we had back then.
The other key question that comes out of this vision is what tools do we need to put into the Portal? To answer this, we need to look carefully at the use cases. I would like to see us break down each of these into a sequence of scientific steps. Then we can ask, which of these tasks can be answered with the Discovery Tool? Where can we call in the Cross-Comparison Tool? And more importantly, what tools are missing? We're also likely to discover that our current tools are not quite tuned to these tasks and need some adjusting.
Actually, having gone through these use cases once, dissecting them can be a bit challenging (but interesting). I did glean, I think, some important and common features that we need to support which I hope to share in my next post. Next week, we'll get more eyes on these use cases, and I hope we can pick out a few exemplary ones which we will use to guide further portal development. So what is the Portal, or rather, what will it become? Ultimately, it is--and I know this is cheap to say--a place to do science. We need to be able tell users what science they can do with the portal; we should be able to describe at least a few of these use cases and have it sound compelling and useful to them.
For my part, at least some of the confusion could be cleared up by answering 'Where is the Portal?" For example, you describe the concept of 'Portal' as this collection of interoperating tools (such as we had in NVO portal of old), however whenever the talk is about software or an actual URL it inevitably comes back to what we know as the Data Discovery tool. So, is DD one part of a 'Portal' website we haven't written yet, or is it the parent project that will consume other functionality (hence the questions about architecture and worries about an all-singing-and-dancing DD tool) with ambitions of monolithic growth?
ReplyDeleteRe Use Cases: I agree these are good use cases and I hope the emphasis on solving them with the DD tool/Portal is being done simply to focus the thinking. These same use cases could be solved in a number of ways, e.g. using other IVOA tools, non-VO tools, cmdline apps, a mixture of all, etc. The question of whether this science can be done, and whether it can be done with the DD-tool/VAO-tools alone, are different. Similarly, there are science-driven use-cases that could be used to focus on DD functionality drivers that aren't on the list. For example, we know that the SN in M101 will create a lot of science in years to come, but trying to find this SN data in the current DD tool exposes issues with finding spectral data, query by time, use of Inventory versus real-time queries, requirements for data aggregators, interactions with the SED tool, and so on. This, I claim, is still science-driven but without having a specific science use-case in mind.
-Mike
> So, is DD one part of a 'Portal' website we haven't written yet
ReplyDeleteYes, that is what I am suggesting.
> These same use cases could be solved in a number of ways, e.g. using other
> IVOA tools, non-VO tools, cmdline apps, a mixture of all, etc.
I certainly agree. I'll note, though, that if one can solve these in any way, we sure as heck better write down how and let our users know!
I do think, though, that it is important that through the portal one can completely solve at least some use cases and solve them easily. Delivering a tool that only takes the user part way there with no other easy way to get to the end of a problem with existing tools is not particularly helpful will not result in much usage.
I'll note also that topics like "SN in M101" come up from time to time: important and obvious topics or problems that you would think we could address but for whatever reason we can't. I think it would be important to get these written down, so that we can consider remedying that situation.
This comment has been removed by the author.
ReplyDeleteIn my opinion, most of the use cases can already be accomplished with the current VO tools, including VAO ones (whew). I would record screencasts and write down tutorials that describe how to accomplish them (also from the command line, using python, for example).
ReplyDeleteHowever, it is likely that some operations are not user friendly or even impossible at the moment, and that's where we would need to focus our attention and efforts.
What I don't understand is why these use cases are in the Portal's scope while the other science initiative's (e.g. SED) use cases are not.
In other words, if the portal is the entry point to the wonderful world of VAO-enabled science, shouldn't I start from the VAO portal in order to get to the SED Tool?
This is not just a matter of semantics or entry points, of course (a list of web links to the download pages of a set of tools doesn't require any design... unless you design an app store!).
Data Discovery requirements, for example, are common to most of the science initiatives, so it makes sense that you use only one data discovery tool, whether you are working with SEDs or Data Mining.
So far I have perceived the Portal has *one of* the VAO tools. What you are proposing sounds different, in that the Portal should glue together all the different tools and provide the scientist with a common framework.
If this is true, then those Science Use Cases are only part of the story, and the portal should inherit all the Science Use Cases that come from the other initiatives, so that you can design the interface between the portal and the tools.
On the other hand, it doesn't look to me that the Portal was designed as such, at least looking at the Project Description Document and requirements:
The new VAO Portal will connect the user to the other key VAO science initiatives
[...]
For year one, this project is expected to support the following other projects, primarily by providing links to their work
In other terms, I agree with Doug, I don't understand what the Portal is. And I thought I did!
If it is what is described in its description document, it is an advanced data discovery tool with links to other projects. Fine. But we need to design the interfaces so that other projects can leverage the portal instead of reimplementing their specific data discovery use cases.
If it is what you are describing here, it will necessarily have to inherit all the science use cases from all the VAO projects and provide an integrated environment that links those together.
The latter vision is great, but as pointed out by Doug it involves the design the VAO architecture and a common framework that all the applications share, not only the Portal. By the way, I though this was the task of the Desktop Integration project which, though, will be a parallel project.
In any case, it looks to me that the Portal wasn't designed with that in mind. If this is the vision for the future of the Portal, it would be useful for it to be designed in a way that would allow the development of plugins (or portlet, as the Portal name would suggest at a first glance).
In that kind of architecture, the SED tool, for example, could have its plugin that would bridge the web based portal to the specific tool (which is desktop based, but something similar would apply to old and new web based services).
The plugin would have its own use cases, derived from the SED use cases and implemented using the common VAO framework, using SAMP under the hood.
10 points from SED for Omar's use of the term "common framework."
ReplyDeletelets just assume that one really wants a monolith common framework/portal to rule them all.
ReplyDeleteNow, lets write out the workflow from data to discovery:
1. find the bits;
2. mange the bits;
3. analyze the bits;
4. visualize the bits;
5. make bytes of new info about the bits.
Take #1. Look at what happens when Datascope hits the registry. After 10+ years of "standardization" a great deal of the work Datascope does as well as many of the bugs/problem that hit the Portal software are due to service variations, poorly managed metadata, etc. Imagine including any *more* service/resource types (spectra, SED). Just collecting the data streams and enabling search against their non-uniform metadata and then aggregating and presenting the results in a roughly uniform fashion so you can move to step #2 is *non-trivial.* This is the complete answer to why the DD portal is where it is today: the science showcase made universal data access its top priority.
so the assertions I'm hearing include one where we should have followed the common framework obsessed PEP and created a portal that provides for all the steps. One that blends into desktop integration. One that utilizes messaging hubs to pass the discovered data to other clients.
Okay but if your port/let vision is roughly akin to Connolly's ASCOT then please be reminded that ASCOT is an example of what can be done when one focuses on a single data feed (SDSS) and doesn't worry about ANY of the heterogeneous data streams that are out there.
let me step back and clarify something: I think that it is totally reasonable to winnow the data feeds down to 1 or a few that make your science cases happiest. The error is in creating use cases that require specific tools or even multiple data providers but not recognizing that these specific tools require us to drop universal access as a principle which limits the data providers that we can hook up. Here are three examples:
The entire discussion about fast image access PRESUMED that one was using a single fast backend to a few important fixed indexable image sets. It wasn't and couldn't be universal (not real time). This would be fine if we negated the premise of universal archive access but we didn't. The effort for building the common framework is still stuck at step #1 while we figure out how to make the fast index stable with the insertion of huge source catalogs, while external catalogs are munged into the index and while we figure out how to augment the indexed results with "real time" ones, which carries a host of curation issues.
take another example: TAP. we can build a TAP based discovery portal around a very few complete (but not yet public) tap implementations but to do so means you HAVE TO throw out pan archive universal data access because universal TAP implementation ain't gonna happen. Worse, we will probably just build the service around 1 maybe 2 TAP services since its clear from the ADS work that your Obscore and my Obscore diverge on what an "observation" means to you or to me.
take a third example: SED. we could build a discovery portal around a few well managed SED data streams. the access interface will be unique against all the other services and we will end up throwing out any that don't comply with SED specs. These issues are likely okay so long as we first negate our premise of universal archive access.
So I guess that is my point about the dream state of "common framework" -- to get to this common framework you have to stop trying to be a common framework and start picking out specialty services to get the framework fleshed out.
Gus,
ReplyDeleteIf I understand your point, I agree that starting using a common framework is just a matter of... well, starting using a common framework ;) And we've got SAMP, and the FASE architecture, and so on.
It also means starting using common policies (e.g. a decentralized orchestration of the samp hub usage, a directory structure in the user's home to keep persistent data, and so on) and sharing reusable components (e.g. a Java extensible framework for SAMP that will be extended by the specific Java projects but whose bulk will be common, etc...).
Is this all in the scope of the Portal? Well, I don't think so, unless the current projects are going to rely, for y2, on a tool which is developed in parallel, and, more importantly, without any definition of a common interface.
I could go on for a lot of time with other (pragmatic) examples... since I apparently share with Doug the impression that we are not dealing with the Portal (whatever it is) but with an Architecture that should be designed *before* coding the projects, and not during the development or even later ;)
On the other hand, I believe that most of your points are related to the registry. The VAO registry should provide access to services that have been strictly validated, so that the single applications are shielded from the problems that arise from non compliant services.
If the data discovery tool provides me with a Spectrum that doesn't even comply to the standards, I have to implement all the sorts of error handling which shouldn't be part of the application (scientific) domain.
ASCOT-like webapps can work if you define a communication framework between portlets and a generic extensible framework for portlets implementation: if you do, you have a decentralized system in which each portlet, implementing only a subset of use cases, offers an interface that can be exploited by other applications.
Then, it is just a matter of layering tasks one on the other, each layer being more complicated (or at a higher level) than the other. Data Discovery is almost on the bottom, using services like the Name Resolver and being using by almost every other tool.
And in this framework you could indeed have "high level" services that use the TAP protocol (via TAP Client API) but with a priori knowledge of the DB schema, so that they can give you access to a set of specific services in a domain consistent fashion (e.g. you know the possible, non standard, values of the classification flag of SDSS and you can offer the user's a selection box with astronomical meaningful labels, star and galaxy, instead of numbers in a SELECT query).
(thanks, guys, for jumping in with the comments!)
ReplyDelete> What I don't understand is why these use cases are in the Portal's scope while
> the other science initiative's (e.g. SED) use cases are not.
> In other words, if the portal is the entry point to the wonderful world of
> VAO-enabled science, shouldn't I start from the VAO portal in order to get to the SED Tool?
Certainly, if you are in the portal there should be a way to leverage Iris, and so yes, it is reasonable to consider the SED use cases in the scope of the "Portal".
(I figured that I might be controversial enough getting people to think about the Portal in terms beyond discovery without dwelling on the role of Iris, ostensibly a desktop tool at the moment, in the midst of a discussion of the Portal, which I have scoped as being a set of browser-based tools. Nevertheless, I had hoped my example of sending discovered SEDs to Iris on the desktop would illustrate its connection to the portal.)
> If it is what is described in its description document, it is an advanced
> data discovery tool with links to other projects. Fine.
> If it is what you are describing here, it will necessarily have to inherit
> all the science use cases from all the VAO projects and provide an integrated
> environment that links those together.
What I am suggesting is closer to the latter than the former, and so, yes, this does brings in a broader collection of use cases. The immediate aim of this exercise is to identify use cases to guide year 2 portal development, so we certainly can't address everything. Thus, we won't have tools for everything right away, we do need think about and document how these tools need to "talk to each other"--the architecture of the framework and a core part of the VAO architecture at large.
I should also say in this context that the tools need to have a loose coupling to avoid a variety of issues one might associate with a "monolith".
> In that kind of architecture, the SED tool, for example, could have its
> plugin that would bridge the web based portal to the specific tool
Beautiful!
> lets just assume that one really wants a monolith common framework/portal to
ReplyDelete> rule them all.
While I recognize this assumption is useful to the discussion Gus presents, I should note that this is not required in practice. I do want users to be able to complete common use cases from start to finish in some fashion. This can be entirely through desktop tools or a combination of portal and desktop (more likely). However, for practical reasons of marketing and software delivery, I believe there should be some broadly applicable use cases that can be solved entirely within the browser environment.
Gus argues that the goal of "universal access" (what I have referred to, I think, as "comprehensive discovery") actually undermines our efforts to build sophisticated research tools that rely on these diverse data sources. I certainly recognize the challenges this poses. It is important to recognize the following realities of the VO architecture:
* Data and services come and go for a variety of reasons
* Some data archives are dynamic: they may grow with time or otherwise provide an infinite number of customized, virtual data products
* Not all archives will support specification X (TAP, ObsTAP, SIA, ...)
* Not all archives will support the same version of specification X
* Not all services will be 100% compliant
Now, I would agree that for the VO to work there is a critical mass of static|up-to-date|compliant services that must available; nevertheless, an architecture that does not allow for these facts is doomed. That is, a discovery architecture that, for example, relies exclusively on universal adoption of ObsTAP at *any mythical time in the future* certainly cannot be sustainable--universal access or no.
The carrot for adoption of and compliance to standards is supposed to be higher visibility. In this way, the NVO Datascope illustrates this fairly well. Putting aside data selection/browsing issues, it didn't matter if 30% of the services failed to respond; Datascope politely explained this and offered access to data it could get. Inevitably there will be somethings that you will want to do that will limit what resources you can access; that is, for the use cases, universal access is not possible.
What are those use cases? That's the purpose of our current use case exploration right now. When I read through the ones assembled in the above mentioned list, I see some that name very specific collections--Chandra, HST, etc. If you are working on a problem like this, is it possible to focus our data discovery to just these tools? There are other use cases that say, "what other data has been collected on these sources?" Clearly, this is a broader question; can we answer it? What we know about the "other data" will depend the level of compliance from the curating archive.
It is my hope that by being use-case driven, we can contend with the realities of the VO architecture more easily. We need not require universal access as required in all behaviors of the tool as long as it is supported where it is needed. For example, fast access does *not* need to available for *all* datasets known to the VO, as long as there is a mechanism to find all datasets for the use cases that call for that.
(I figured that I might be controversial enough getting people to think about the Portal in terms beyond discovery without dwelling on the role of Iris, ostensibly a desktop tool at the moment, in the midst of a discussion of the Portal, which I have scoped as being a set of browser-based tools. Nevertheless, I had hoped my example of sending discovered SEDs to Iris on the desktop would illustrate its connection to the portal.)
ReplyDeleteRay, this means that my point wasn't clear, and I apologize for that, especially after such a long comment ;)
My point is that "sending a discovered SED to Iris" is just one of a broad range of use cases in a "common framework", i.e. an environment that adds value to the single tools by allowing the user to employ different tools in a creative way.
My view of an effort like the VAO (as opposed to the IVOA in general) is that you can add value by enclosing the single tools (even third party tools like Aladin) in a more consistent and focused environment.
"Sending something to sometool" is only part of this, the part on which you rely to build layers on the top of each other.
For example, you can extend the object oriented event paradigm in a multi-language, multi-process fashion by registering listeners from one tool to the other, using SAMP as a protocol for accomplishing this... in this example you are not sending SEDs, but event objects serialized as SAMP messages. Notice that this is how, for example, Sherpa throws exceptions from Python that are caught by Specview in Java, without the user even realizing the complex system underneath Iris.
I thought that with this new vision of the Portal you intended something like this, because you mentioned the Portal as a "single entry point" to the VAO tools.
In any case, keeping the example of Iris, for year two we have several requirements that massively overlap with the Data Discovery Tool: they derive from (or can lead to) use cases that are not included in the Portal scientific use cases. Neither the portal seems to make room for an Iris Plugin that would bridge the portal to Iris and implement the Iris specific use cases (this is not the only option, just an example).
Without an architectural view and an interaction/interface among projects (both in the human and in the application space) we risk to replicate work and/or to miss important requirements, imho.
Thus, with my initial comment I wanted to point out that if those Use Cases make it to the Portal's plate and the others don't we might miss the opportunity to integrate the different projects in a "common framework", whatever the Portal is.
In other terms, the Portal would satisfy its Use Cases subset, Iris the others, but being two monoliths on their own, instead of two components in the same dynamic environment, even if you can beam files between each other.
Even more importantly, we currently lack a general documentation framework, which would be another important use case in the scope of a "single entry point" and would make my point even stronger.
I would expect that the Portal could "answer" my generic question: "how do I analyze SEDs"? or it could allow me to use keywords, like sed, spectrum, cross comparison, and provide me with a set of resources, tutorials, screencasts, download links that help me getting my science done. This is not just a page full of links or an app store.
In this case, all the projects would coherently contribute to the main (documentation) framework with their modules, and would point to each other... for example Iris might point the user to the Comparison Tool, even though Cross Comparison is not in the Iris scope.
Coherence, I believe that's my point: without coherence, with 20W you can turn a light bulb on. Useful. But with coherence and the same 20W you can power a laser beam used for cutting microprocessors. More useful.
Sorry, I didn't notice your update... well, all my comments apply to any combination of web/desktop tools and even to a (strictly defined) portal that enclose different modules satisfying different use cases.
ReplyDeleteIn other terms they apply to any "entryway to a collection of tools that work together".