Your SlideShare is downloading. ×
DataUp Presentation at Cal Poly
Upcoming SlideShare
Loading in...5
×

Thanks for flagging this SlideShare!

Oops! An error has occurred.

×

Introducing the official SlideShare app

Stunning, full-screen experience for iPhone and Android

Text the download link to your phone

Standard text messaging rates apply

DataUp Presentation at Cal Poly

365
views

Published on

Presentation on the DataUp Tool at the Cal Poly Kennedy Library, 18 April 2013.

Presentation on the DataUp Tool at the Cal Poly Kennedy Library, 18 April 2013.

Published in: Sports, Technology

0 Comments
1 Like
Statistics
Notes
  • Be the first to comment

No Downloads
Views
Total Views
365
On Slideshare
0
From Embeds
0
Number of Embeds
1
Actions
Shares
0
Downloads
3
Comments
0
Likes
1
Embeds 0
No embeds

Report content
Flagged as inappropriate Flag as inappropriate
Flag as inappropriate

Select your reason for flagging this presentation as inappropriate.

Cancel
No notes for slide

Transcript

  • 1. Carly  Strasser    California  Digital  Library    @carlystrasser  April  2013  @  DataUp:    Helping  manage  &  archive  data    From  Flickr  by  Spatial  Mongrel  
  • 2. C.  Strasser  C.  Strasser  C.  Strasser  C.  Strasser  Courtesy  of  WHOI  
  • 3. Why  don’t  people  share  data?  Is  data  management  being  taught?  Do  attitudes  about  sharing  differ  among  disciplines?  What  role  can  libraries  play  in  data  education?  How  can  we  promote  storing  data  in  repositories?  What  barriers  to  sharing  can  we  eliminate?  
  • 4. Why  is  data  management      a  hot  topic?  From  Flickr  by  Velo  Steve  
  • 5. Back in the day…Da  Vinci  Curie  Newton  classicalschool.blogspot.com  Darwin  
  • 6. Digital  data  From  Flickr  by  Flickmor  From  Flickr  by  US  Army  Environmental  Command  From  Flickr  by    DW0825  C.  Strasser  Courtesey  of  WHOI  From  Flickr  by    deltaMike  
  • 7. Digital  data  +    Complex  workflows  
  • 8. C:Documents and SettingshamptonMy DocumentsNCEAS Distributed Graduate Seminars[Wash Cres Lake Dec 15 Dont_Use.xls]Sheet1Stable Isotope Data SheetWash Cresc Lake Peters lab Dont use - old dataAlgal Washed RocksDec. 16Tray 004SD for delta13C = 0.07 SD for delta15N = 0.15Position SampleID Weight (mg) %C delta 13C delta 13C_ca %N delta 15N delta 15N_ca Spec. No.A1 ref 0.98 38.27 -25.05 -24.59 1.96 4.12 3.47 25354A2 ref 0.98 39.78 -25.00 -24.54 2.03 4.01 3.36 25356A3 ref 0.98 40.37 -24.99 -24.53 2.04 4.09 3.44 25358A4 ref 1.01 42.23 -25.06 -24.60 2.17 4.20 3.55 25360 Shore Avg ConA5 ALG01 3.05 1.88 -24.34 -23.88 0.17 -1.65 -2.30 25362 c -1.26 -27.22A6 Lk Outlet Alg 3.06 31.55 -30.17 -29.71 0.92 0.87 0.22 25364 1.26 0.32A7 ALG03 2.91 6.85 -21.11 -20.65 0.48 -0.97 -1.62 25366 cA8 ALG05 2.91 35.56 -28.05 -27.59 2.30 0.59 -0.06 25368A9 ALG07 3.04 33.49 -29.56 -29.10 1.68 0.79 0.14 25370A10 ALG06 2.95 41.17 -27.32 -26.86 1.97 2.71 2.06 25372B1 ALG04 3.01 43.74 -27.50 -27.04 1.36 0.99 0.34 25374 cB2 ALG02 3 4.51 -22.68 -22.22 0.34 4.31 3.66 25376B3 ALG01 2.99 1.59 -24.58 -24.12 0.15 -1.69 -2.34 25378 cB4 ALG03 2.92 4.37 -21.06 -20.60 0.34 -1.52 -2.17 25380 cB5 ALG07 2.9 33.58 -29.44 -28.98 1.74 0.62 -0.03 25382B6 ref 1.01 44.94 -25.00 -24.54 2.59 3.96 3.31 25384B7 ref 0.99 42.28 -24.87 -24.41 2.37 4.33 3.68 25386B8 Lk Outlet Alg 3.04 31.43 -29.69 -29.23 1.07 0.95 0.30 25388B9 ALG06 3.09 35.57 -27.26 -26.80 1.96 2.79 2.14 25390B10 ALG02 3.05 5.52 -22.31 -21.85 0.45 4.72 4.07 25392C1 ALG04 2.98 37.90 -27.42 -26.96 1.36 1.21 0.56 25394 cC2 ALG05 3.04 31.74 -27.93 -27.47 2.40 0.73 0.08 25396C3 ref 0.99 38.46 -25.09 -24.63 2.40 4.37 3.72 2539823.78 1.17Reference statistics:Sampling Site / Identifier:Sample Type:Date:Tray ID and Sequence:From  Stephanie  Hampton  (2010)      ESA  Workshop  on  Best  Practices  2  tables   Random  notes  From  Stephanie  Hampton  
  • 9. C:Documents and SettingshamptonMy DocumentsNCEAS Distributed Graduate Seminars[Wash Cres Lake Dec 15 Dont_Use.xls]Sheet1Stable Isotope Data SheetWash Cresc Lake Peters lab Dont use - old dataAlgal Washed RocksDec. 16Tray 004SD for delta13C = 0.07 SD for delta15N = 0.15Position SampleID Weight (mg) %C delta 13C delta 13C_ca %N delta 15N delta 15N_ca Spec. No.A1 ref 0.98 38.27 -25.05 -24.59 1.96 4.12 3.47 25354A2 ref 0.98 39.78 -25.00 -24.54 2.03 4.01 3.36 25356A3 ref 0.98 40.37 -24.99 -24.53 2.04 4.09 3.44 25358A4 ref 1.01 42.23 -25.06 -24.60 2.17 4.20 3.55 25360 Shore Avg ConA5 ALG01 3.05 1.88 -24.34 -23.88 0.17 -1.65 -2.30 25362 c -1.26 -27.22A6 Lk Outlet Alg 3.06 31.55 -30.17 -29.71 0.92 0.87 0.22 25364 1.26 0.32A7 ALG03 2.91 6.85 -21.11 -20.65 0.48 -0.97 -1.62 25366 cA8 ALG05 2.91 35.56 -28.05 -27.59 2.30 0.59 -0.06 25368A9 ALG07 3.04 33.49 -29.56 -29.10 1.68 0.79 0.14 25370A10 ALG06 2.95 41.17 -27.32 -26.86 1.97 2.71 2.06 25372B1 ALG04 3.01 43.74 -27.50 -27.04 1.36 0.99 0.34 25374 cB2 ALG02 3 4.51 -22.68 -22.22 0.34 4.31 3.66 25376B3 ALG01 2.99 1.59 -24.58 -24.12 0.15 -1.69 -2.34 25378 cB4 ALG03 2.92 4.37 -21.06 -20.60 0.34 -1.52 -2.17 25380 cB5 ALG07 2.9 33.58 -29.44 -28.98 1.74 0.62 -0.03 25382B6 ref 1.01 44.94 -25.00 -24.54 2.59 3.96 3.31 25384B7 ref 0.99 42.28 -24.87 -24.41 2.37 4.33 3.68 25386B8 Lk Outlet Alg 3.04 31.43 -29.69 -29.23 1.07 0.95 0.30 25388B9 ALG06 3.09 35.57 -27.26 -26.80 1.96 2.79 2.14 25390B10 ALG02 3.05 5.52 -22.31 -21.85 0.45 4.72 4.07 25392C1 ALG04 2.98 37.90 -27.42 -26.96 1.36 1.21 0.56 25394 cC2 ALG05 3.04 31.74 -27.93 -27.47 2.40 0.73 0.08 25396C3 ref 0.99 38.46 -25.09 -24.63 2.40 4.37 3.72 2539823.78 1.17Reference statistics:Sampling Site / Identifier:Sample Type:Date:Tray ID and Sequence:From  Stephanie  Hampton  (2010)      ESA  Workshop  on  Best  Practices  Wash  Cres  Lake  Dec  15  Dont_Use.xls  From  Stephanie  Hampton  
  • 10. C:Documents and SettingshamptonMy DocumentsNCEAS Distributed Graduate Seminars[Wash Cres Lake Dec 15 Dont_Use.xls]Sheet1Stable Isotope Data SheetWash Cresc Lake Peters lab Dont use - old dataAlgal Washed RocksDec. 16Tray 004SD for delta13C = 0.07 SD for delta15N = 0.15Position SampleID Weight (mg) %C delta 13C delta 13C_ca %N delta 15N delta 15N_ca Spec. No.A1 ref 0.98 38.27 -25.05 -24.59 1.96 4.12 3.47 25354A2 ref 0.98 39.78 -25.00 -24.54 2.03 4.01 3.36 25356A3 ref 0.98 40.37 -24.99 -24.53 2.04 4.09 3.44 25358A4 ref 1.01 42.23 -25.06 -24.60 2.17 4.20 3.55 25360 Shore Avg ConA5 ALG01 3.05 1.88 -24.34 -23.88 0.17 -1.65 -2.30 25362 c -1.26 -27.22A6 Lk Outlet Alg 3.06 31.55 -30.17 -29.71 0.92 0.87 0.22 25364 1.26 0.32A7 ALG03 2.91 6.85 -21.11 -20.65 0.48 -0.97 -1.62 25366 cA8 ALG05 2.91 35.56 -28.05 -27.59 2.30 0.59 -0.06 25368A9 ALG07 3.04 33.49 -29.56 -29.10 1.68 0.79 0.14 25370A10 ALG06 2.95 41.17 -27.32 -26.86 1.97 2.71 2.06 25372B1 ALG04 3.01 43.74 -27.50 -27.04 1.36 0.99 0.34 25374 c SUMMARY OUTPUTB2 ALG02 3 4.51 -22.68 -22.22 0.34 4.31 3.66 25376B3 ALG01 2.99 1.59 -24.58 -24.12 0.15 -1.69 -2.34 25378 c Regression StatisticsB4 ALG03 2.92 4.37 -21.06 -20.60 0.34 -1.52 -2.17 25380 c Multiple R 0.283158B5 ALG07 2.9 33.58 -29.44 -28.98 1.74 0.62 -0.03 25382 R Square 0.080178B6 ref 1.01 44.94 -25.00 -24.54 2.59 3.96 3.31 25384 Adjusted R Square-0.022024B7 ref 0.99 42.28 -24.87 -24.41 2.37 4.33 3.68 25386 Standard Error1.906378B8 Lk Outlet Alg 3.04 31.43 -29.69 -29.23 1.07 0.95 0.30 25388 Observations 11B9 ALG06 3.09 35.57 -27.26 -26.80 1.96 2.79 2.14 25390B10 ALG02 3.05 5.52 -22.31 -21.85 0.45 4.72 4.07 25392 ANOVAC1 ALG04 2.98 37.90 -27.42 -26.96 1.36 1.21 0.56 25394 c df SS MS F Significance FC2 ALG05 3.04 31.74 -27.93 -27.47 2.40 0.73 0.08 25396 Regression 1 2.851116 2.851116 0.784507 0.398813C3 ref 0.99 38.46 -25.09 -24.63 2.40 4.37 3.72 25398 Residual 9 32.7085 3.63427823.78 1.17 Total 10 35.55962CoefficientsStandard Error t Stat P-value Lower 95%Upper 95%Lower 95.0%Upper 95.0%Intercept -4.297428 4.671099 -0.920003 0.381568 -14.8642 6.269341 -14.8642 6.269341X Variable 1-0.158022 0.17841 -0.885724 0.398813 -0.561612 0.245569 -0.561612 0.245569Reference statistics:Sampling Site / Identifier:Sample Type:Date:Tray ID and Sequence:Random  stats  output  From  Stephanie  Hampton  
  • 11. C:Documents and SettingshamptonMy DocumentsNCEAS Distributed Graduate Seminars[Wash Cres Lake Dec 15 Dont_Use.xls]Sheet1Stable Isotope Data SheetWash Cresc Lake Peters lab Dont use - old dataAlgal Washed RocksDec. 16Tray 004SD for delta13C = 0.07 SD for delta15N = 0.15Position SampleID Weight (mg) %C delta 13C delta 13C_ca %N delta 15N delta 15N_ca Spec. No.A1 ref 0.98 38.27 -25.05 -24.59 1.96 4.12 3.47 25354A2 ref 0.98 39.78 -25.00 -24.54 2.03 4.01 3.36 25356A3 ref 0.98 40.37 -24.99 -24.53 2.04 4.09 3.44 25358A4 ref 1.01 42.23 -25.06 -24.60 2.17 4.20 3.55 25360 Shore Avg ConA5 ALG01 3.05 1.88 -24.34 -23.88 0.17 -1.65 -2.30 25362 c -1.26 -27.22A6 Lk Outlet Alg 3.06 31.55 -30.17 -29.71 0.92 0.87 0.22 25364 1.26 0.32A7 ALG03 2.91 6.85 -21.11 -20.65 0.48 -0.97 -1.62 25366 cA8 ALG05 2.91 35.56 -28.05 -27.59 2.30 0.59 -0.06 25368A9 ALG07 3.04 33.49 -29.56 -29.10 1.68 0.79 0.14 25370A10 ALG06 2.95 41.17 -27.32 -26.86 1.97 2.71 2.06 25372B1 ALG04 3.01 43.74 -27.50 -27.04 1.36 0.99 0.34 25374 c SUMMARY OUTPUTB2 ALG02 3 4.51 -22.68 -22.22 0.34 4.31 3.66 25376B3 ALG01 2.99 1.59 -24.58 -24.12 0.15 -1.69 -2.34 25378 c Regression StatisticsB4 ALG03 2.92 4.37 -21.06 -20.60 0.34 -1.52 -2.17 25380 c Multiple R 0.283158B5 ALG07 2.9 33.58 -29.44 -28.98 1.74 0.62 -0.03 25382 R Square 0.080178B6 ref 1.01 44.94 -25.00 -24.54 2.59 3.96 3.31 25384 Adjusted R Square-0.022024B7 ref 0.99 42.28 -24.87 -24.41 2.37 4.33 3.68 25386 Standard Error1.906378B8 Lk Outlet Alg 3.04 31.43 -29.69 -29.23 1.07 0.95 0.30 25388 Observations 11B9 ALG06 3.09 35.57 -27.26 -26.80 1.96 2.79 2.14 25390B10 ALG02 3.05 5.52 -22.31 -21.85 0.45 4.72 4.07 25392 ANOVAC1 ALG04 2.98 37.90 -27.42 -26.96 1.36 1.21 0.56 25394 c df SS MS F Significance FC2 ALG05 3.04 31.74 -27.93 -27.47 2.40 0.73 0.08 25396 Regression 1 2.851116 2.851116 0.784507 0.398813C3 ref 0.99 38.46 -25.09 -24.63 2.40 4.37 3.72 25398 Residual 9 32.7085 3.63427823.78 1.17 Total 10 35.55962CoefficientsStandard Error t Stat P-value Lower 95%Upper 95%Lower 95.0%Upper 95.0%Intercept -4.297428 4.671099 -0.920003 0.381568 -14.8642 6.269341 -14.8642 6.269341X Variable 1-0.158022 0.17841 -0.885724 0.398813 -0.561612 0.245569 -0.561612 0.245569Reference statistics:Sampling Site / Identifier:Sample Type:Date:Tray ID and Sequence:SampleID ALG03 ALG05 ALG07 ALG06 ALG04 ALG02 ALG01 ALG03 ALG07Weight (mg) 2.91 2.91 3.04 2.95 3.01 3 2.99 2.92 2.9%C 6.85 35.56 33.49 41.17 43.74 4.51 1.59 4.37 33.58delta 13C -21.11 -28.05 -29.56 -27.32 -27.50 -22.68 -24.58 -21.06 -29.44delta 13C_ca -20.65 -27.59 -29.10 -26.86 -27.04 -22.22 -24.12 -20.60 -28.98%N 0.48 2.30 1.68 1.97 1.36 0.34 0.15 0.34 1.74delta 15N -0.97 0.59 0.79 2.71 0.99 4.31 -1.69 -1.52 0.62delta 15N_ca -1.62 -0.06 0.14 2.06 0.34 3.66 -2.34 -2.17 -0.03-3.00-2.00-1.000.001.002.003.004.00-35.00 -30.00 -25.00 -20.00 -15.00 -10.00 -5.00 0.00Series1From  Stephanie  Hampton  
  • 12. UGLY TRUTHData  management?  Metadata?  Data  repositories?  Share  data  publicly?  Why  share  data?    From  Flickr  by  s  i  b  e  r  ABOUTRESEARCHERS
  • 13. Who  cares?  From  Flickr  by  Redden-­‐McAllister  From  Flickr  by  AJC1  
  • 14. From  Flickr  by  Michael  Tinkler  ?  
  • 15. From  Flickr  by  iowa_spirit_walker  •  Cost  •  Confusion  about  standards  •  Lack  of  training  •  Fear  of  lost  rights  or  benefits  •  No  incentives  
  • 16. From  Flickr  by  thewma1  
  • 17. Intercept  researchers  where  they  already  work  
  • 18. Facilitate  Archiving  Sharing  Publishing  Data  management  &  organization  Data  Reuse  &  Reproducibility  
  • 19. What  do  scientists  need  help  with?  
  • 20. Asked  ~200  scientists  What  does  your  data  look  like?  How  do  you  capture  metadata?  Plans  for  saving  &  sharing  data?  Repositories?  
  • 21. What  the  tool  should  do:    Best  practices  check  Generate  metadata  (EML)  Get  identifier  +  citation  Post  data  to  repository    From  Flickr  by  Rennett  Stowe  
  • 22. Open  Source  Tool   Add-­‐in  &  Web  Application  csv  &  xlsx  dataup.cdlib.org  Free  ?
  • 23. Add-­‐in    •  Software  you  download  &  install  •  Appears  as  “ribbon”  in  Excel  •  Works  for  Windows  Excel  2007+  Web-­‐based  application    •  Upload  file  to  website  •  Works  for  any  platform  •  But…  new  user  interface  VS  
  • 24. DataUp    Features    Best  practices  check  Generate  metadata  Get  identifier  &  citation  Post  data  to  repository  From  Flickr  by  SoulRider.222  
  • 25. Best  Practices  Check  •  Embedded  charts,  tables,  pictures  •  Embedded  comments  •  Commas  •  Special  characters  •  Color-­‐coded  text  &  cell  shading  •  Columns  with  mixed  data  types  •  Non-­‐contiguous  data  •  Merged  cells  •  Blank  cells  •  No  header  row    •  Multiple  sheets  From  Flickr  by  ex.libris  
  • 26. DataUp    Features    Best  practices  check  Generate  metadata  Get  identifier  &  citation  Post  data  to  repository  From  Flickr  by  SoulRider.222  
  • 27. •  Digital  context  •  Name  of  the  data  set  •  The  name(s)  of  the  data  file(s)  in  the  data  set  •  Date  the  data  set  was  last  modified  •  Example  data  file  records  for  each  data  type  file  •  Pertinent  companion  files  •  List  of  related  or  ancillary  data  sets  •  Software  (including  version  number)  used  to  prepare/read    the  data  set  •  Data  processing  that  was  performed  •  Personnel  &  stakeholders  •  Who  collected    •  Who  to  contact  with  questions  •  Funders  •  Scientific  context  •  Scientific  reason  why  the  data  were  collected  •  What  data  were  collected  •  What  instruments  (including  model  &  serial  number)  were  used  •  Environmental  conditions  during  collection  •  Where  collected  &  spatial  resolution  When  collected  &  temporal  resolution  •  Standards  or  calibrations  used  •  Information  about  parameters  •  How  each  was  measured  or  produced  •  Units  of  measure  •  Format  used  in  the  data  set  •  Precision  &  accuracy  if  known  •  Information  about  data  •  Definitions  of  codes  used  •  Quality  assurance  &  control  measures  •  Known  problems  that  limit  data  use  (e.g.  uncertainty,  sampling  problems)    •  How  to  cite  the  data  set  Holy  Metadata!  
  • 28. ~45  elements  included    7  required  Creator  details  Title  Date  Keywords  Abstract    File-­‐level  Metadata  
  • 29. Name  Definition  Type  (text,  date/time,  numeric)  Unit  Location  (sheet)  Attribute  Metadata  
  • 30. DataUp    Features    Best  practices  check  Generate  metadata  Get  identifier  &  citation  Post  data  to  repository  From  Flickr  by  SoulRider.222  
  • 31. Identifier  +  Citation  Allows  readers  to  find  data  products  Get  credit  for  data  and  publications  Promotes  reproducibility  Better  measure  of  research  impact  Example:  Sidlauskas,  B.  2007.  Data  from:  Testing  for  unequal  rates  of  morphological  diversification  in  the  absence  of  a  detailed  phylogeny:  a  case  study  from  characiform  fishes.  Dryad  Digital  Repository.  doi:10.5061/dryad.20   Persistent  Unique  Identifier  From  Flickr  by  maybeemily  
  • 32. DataUp    Features    Best  practices  check  Generate  metadata  Get  identifier  &  citation  Post  data  to  repository  From  Flickr  by  SoulRider.222  
  • 33. Data  Repository  for  Anyone  |  Anywhere  
  • 34. DataUp  Web  App  
  • 35. Web  App  
  • 36. Web  App  
  • 37. Web  App:  Best  Practices  Check  
  • 38. Web  App:  Metadata  
  • 39. Web  App:  Citation  
  • 40. Web  App:  Posting  to  repository  
  • 41. DataUp  Add-­‐In  
  • 42. Add-­‐in:  Ribbon  
  • 43. Add-­‐in:  Metadata  tab  
  • 44. Add-­‐In  •  Windows  PC  2007+  •  No  log-­‐in  required  •  Offline  &  online  •  Can  view  metadata  via  tab  •  See  check  alongside  data  •  Select  header  row  Web  app  •  Any  platform  •  Log-­‐in  required  •  Online  only  •  Can’t  view  metadata  once  generated  •  Get  locations  for  check  •  Manual  header  row  entry  VS  
  • 45. Main  site:  dataup.cdlib.org  
  • 46. From  Flickr  by  Hanna-­‐  •  No  knowledge  of  code  •  No  C#/.NET    •  No  Visual  Basic    •  No  money  
  • 47. From  Flickr  by  401(K)  2013  
  • 48. •  New  language  •  Focus  on  web  app  •  Emphasize  Best  Practices  Check  •  Leverage  existing  tools  •  Enable  Customization  From  animationresources.org  
  • 49. dataup.cdlib.org  bitbucket.org/dataup/main  Website  Code  site  My  website  Email  me  Tweet  me  My  slides  CDL  Blog  carlystrasser.net  carlystrasser@gmail.com  @carlystrasser    slideshare.net/carlystrasser  datapub.cdlib.org