Friday, September 9, 2011

USVAO Web site overview

With the imminent delivery of VAO science services, the organization of the VAO web site is an urgent issue.  We are going to make decisions in the next few weeks as to where users will see the VAO portal, SED, cross-correlation and time series services.  Recently Sarah asked if we should create a new iris.usvao.org virtual host to provide a location for us to serve the IRIS service.  My suggestion was to wait a few days so that we might think about this in the context of what we want the VAO web site to look like.

First my quick summary of where we currently stand.  This is intended to list the web sites that people are charging the VAO to support.  I'm sure there are inaccuracies.

Services currently using USVAO web address:
  • www.usvao.org    Primary Web site (at CalTech) 
  • help.usvao.org     JIRA (at NOAO)
  • dev.usvao.org      TRAC (at NCSA?)
  • wiki.usvao.org     Wiki (at CalTech)
  • portal.usvao.org  Portal GUI (at ST ScI)

VAO supported services not using USVAO Web address:
  • nvo.stsci.edu/vor10                            VAO registry
  • cxc.cfa.harvard.edu/csc1/temp/sed    IRIS help and download
  • heasarc.gsfc.nasa.gov/vo/...               Validation, monitoring and notification
  • voservices.net/nvolog                         Logging
  • nvo.ncsa.uiuc.edu/dalvalidate/...        Validators
  • iraf-nvo.noao.edu/vo-cli                     VO Client services (site?)
  • sso.us-vo.org                                       Single Sign on service (NCSA/NOAO)
  • astrobabel.com                                   Forum
Services supported on social media web sites.
  • usvao.blogspot.com    VAO Blog
  • twitter.com/usvao       Twitter
  • facebook.com/usvao    Facebook
  • calendar.google.com    Calendars 
External services supported
  • ivoa.net           IVOA website/reg. of registries (rofr.ivoa.net)
  • skyalert.org     Not sure this belongs or perhaps its moving to  being a core function...
Not yet released (I don't know the full URLs here)
  • ipac.caltech.edu Cross-correlation service, time series service
  • jhu.edu               Cross-correlation service, TAP server
  • cfa.harvard.edu  TAP Client
  • stsci.edu              EPO pages
This does not include services that are intended primarily for non-interactive consumption (e.g., the inventory and DataScope web services used in the Data Discovery Service).  These are the addresses that a user trying to find/use a VAO supported capability might need to know about.

My belief is that simply adding new virtual web sites  willy nilly like
   iris.usvao.org
and
  timeseries.usvao.org
will lead to a cluttered and potentially confusing web site. In fact I think we're already getting there.

So here's a suggested strawman organization for the VAO web site.  I'm not particularly wed to the specific
tags or this structure but I do think that having some structure will make it easier for us and our users to follow.


www.usvao.org                           (current)
    science.usvao.org                     (new -- directed towards scientists)
       /discovery       (current portal.usvao.org)
       /iris                 (currently at Harvard)
       /tapexplorer    (TAP Client)
       /xcorr             (integrate TBR x-corr interfaces)
       /time               (TBR timeseries services)

   support.usvao.org                     (new -- directed towards developers/institutions)
       /software                               (new)
          /clientTools                         (Voclient and such)
          /serverTools                       (DALServer and TAPserver)
          /svn                                   (Link to SVN or its replacement)
       /docs                                    (new probably has lots of custom links underneath)
          /repository                         (documentation repository)
          /wiki                                  (wiki if visible to public)
          /devel                                (trac if visible to public)
       /status                                   (operations validation/stuff currently at HEASARC)

     help.usvao.org                     (reuse name.  Pages that help users.)
          /request                             (Form for users to submit requests to user input)
          /internal                             (the current top level page. JIRA for VAO users only)
          /forum                               (the current forum)
          /staff                                  (pages to help users identify/contact staff)

    epo.usvao.org                          (new, directed towards teachers, students, public)
  
    internal.usvao.org                     (new, directed to VAO staff)
        /resources                            Resources that must be hosted outside our web site
        /staffGuide                           Documentation to help employees understand VAO.    
        /dev                                     Non-public development pages (current dev.usavo.org)
        /ops                                     Non-public operations pages
        /logs                                     Logging        


Please feel free to comment.  We may try to discuss this at the Ops telecon next Monday but this clearly affects all of us not just ops.

One technical issue does need to be kept in mind.  Given a remote site it is possible to link to it as both

      remote.usvao.org
or
      usvao.org/remote

but the mechanisms are different and we may have access to only one or the other.  So at least in the short term we may not always be able to get the site we want.  However I think we should first design the site we'd want to see and then recognize that we may need to have deviations.

Saturday, September 3, 2011

What is the Portal?

As you might know, many of us are meeting next week in Baltimore (physically and virtually) to discuss portal development in Year 2 (which we plan to dive into right after the second beta release of the Data Discovery Tool at the end of September). This meeting is part of an effort across all our science initiatives to ensure that our development is science-driven. Consequently, the big focus of this meeting will be science use cases. Gretchen Greene has assembled a great set of use cases that were developed by a number of people reaching back to the NVO days and spanning to more recent examples from the VAO and the IVOA. This week I spent some time with these use cases to try to pull out some common features and needs. It was an inspiring and worthwhile exercise--I highly recommend it.

Doug Tody made the comment this week that considering these use cases as they are laid out on that page makes this meeting more of one about the VAO architecture. I believe he was making reference to the fact that our first year's effort for the portal was about data discovery, resulting in a nicely focused tool that includes some new techniques for quickly finding and selecting datasets. The use cases, on the other hand, go beyond discovery. So, yes, I have to agree with him: addressing these use cases does mean addressing the VAO architecture.

He also said in the same breath, "I'm not sure what the Portal is." As we take a step back to consider what we've accomplished in the first year and what we'll aim for in the second, it's a good time to try re-answering that question. We should do it in terms of science use cases--what it is we want astronomers to accomplish--and, yes, this means addressing some of the essentials of the VAO architecture.

I want to suggest that we see the Portal as a web-based entryway to a collection of tools that work together. What you can do in the portal goes beyond data discovery; that is, the Data Discovery Tool is just one of those tools. The Cross Comparison Tool should be another (actually quite a critical one, as I hope to discuss in a later post). The TAP client should have a presence there as well. And what can you accomplish in the portal? Those use cases, of course.

Updated 9/6: Given the discussion below, I should note that I meant to say that the Portal--in my definition--is explicitly browser-based (with programmatic web service interfaces possible, too).

In this vision, the Portal is not a monolithic tool that does everything; each tool specializes in some task, and over time we can add more tools. It's worth noting though that like all good use cases, ours say nothing about "tools." Clearly, to solve those problems, one will have to use several tools together. Thus, a key aim of the Portal is allowing the tools to work together: the user should be able to easily discover and create lists of datasets and sources with the Discovery Tool and join them together with the Cross-Comparison Tool. It would be great if that collaboration of tools were so quick and smooth, that the user may not realize that two different tools were used. Of course, this kind of collaboration extends to the desktop: the user should be able to send selected catalog records to Topcat and a list of SEDs to Iris.

A key question, then, we need to answer in the design of the Portal is what is needed to allow these tools to work together? This, some of you will recall, was a key problem that had to be tackled when we were developing the NVO portal. Fortunately, we have a few more technological options available to us today than we had back then.

The other key question that comes out of this vision is what tools do we need to put into the Portal? To answer this, we need to look carefully at the use cases. I would like to see us break down each of these into a sequence of scientific steps. Then we can ask, which of these tasks can be answered with the Discovery Tool? Where can we call in the Cross-Comparison Tool? And more importantly, what tools are missing? We're also likely to discover that our current tools are not quite tuned to these tasks and need some adjusting.

Actually, having gone through these use cases once, dissecting them can be a bit challenging (but interesting). I did glean, I think, some important and common features that we need to support which I hope to share in my next post. Next week, we'll get more eyes on these use cases, and I hope we can pick out a few exemplary ones which we will use to guide further portal development. So what is the Portal, or rather, what will it become? Ultimately, it is--and I know this is cheap to say--a place to do science. We need to be able tell users what science they can do with the portal; we should be able to describe at least a few of these use cases and have it sound compelling and useful to them.