hq workshop 2014
DESCRIPTION
HQ Workshop 2014. OTN SandBox Presented by Marta Mihoff OTN Database/Data Process Manager. Start Oracle Virtual Box and OTN SandBox. Open Start Window Click on Oracle VM VirtualBox. Start OTN Sandbox. Outline Background Platform Overview Data and Program Management Exercises - PowerPoint PPT PresentationTRANSCRIPT
![Page 1: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/1.jpg)
HQ Workshop 2014OTN SandBox
Presented by Marta Mihoff OTN Database/Data Process Manager
![Page 2: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/2.jpg)
Start Oracle Virtual Box and OTN SandBox
• Open Start Window • Click on Oracle VM VirtualBox
• Start OTN Sandbox
![Page 3: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/3.jpg)
Outline• Background• Platform Overview• Data and Program Management• Exercises• Update the ‘sandbox’ folder• Create working folder• Data folder management• File Conversion• White-Mihoff False Filtering Tool• Distance Matrix Merge• Mihoff Interval Data Tool• Cleanup Tool
• Wrap Up
![Page 4: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/4.jpg)
OTN Sandbox Backround• Symposium 2013 researcher requests• Fall 2013 request from Steve Kessel and
Eddie Halfyard • Reverse engineered Easton White’s code• Presentation Platform• Winter development and testing
![Page 5: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/5.jpg)
OTN SandBox Platform
• Free open software Black Box•Oracle Virtual Box•Vagrant VMware Integration
![Page 6: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/6.jpg)
•OTN Sandbox Appliance
• Postgresql 9.1 database• PGAdmin• Python 2.7 •Rstudio (only part visible)• TMB – Statistical Modelling Package (R)
![Page 7: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/7.jpg)
OTN SandBox Tools• White-Mihoff False Filtering Tool• Builds a file of suspect detections• Creates a file of filtered detections• Creates a distance matrix
• Distance Matrix Merge• Outputs a matrix overriding distances with researcher input
• Mihoff Interval Data Tool• Creates a file of Compressed detections and a file of interval data
• Miscellaneous• File Conversion (UTF8)• Cleanup
![Page 8: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/8.jpg)
Start Sandbox• Click Start button
• Click cmd.exe to open command box
• Navigate to your OTNsandbox folder (see install document)
• Then execute command ‘vagrant up’
• Type cmd in search box
![Page 9: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/9.jpg)
Update OTN sandbox folder
• Click the Icon on your task bar
• Click the Icon on L.H.S.
![Page 10: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/10.jpg)
Update OTN sandbox folder
• Type ‘cd RStudio/sandbox’ on the command line
• Enter
• Type ‘git pull’ on the command line
• Enter
![Page 11: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/11.jpg)
Sign In
• Open Chrome or Firefox • Paste sandbox URL
• Sign in• Username: sandbox• Password: otn123
• Will not work with VPN turned on
![Page 12: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/12.jpg)
Data Folder Management
• Manage your own data• Current working data folder is always “data”• The “data” folder in Rstudio is a direct link to folder
“data” in your Desktop/OTNsandbox folder 1. Never Delete or Rename folder “data”2. You can copy folder “data”3. And you can copy into folder “data”
![Page 13: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/13.jpg)
Program Folder Management
• All programs are in folder `sandbox`• Warning:• This folder contains all the software developed by
the OTN team.• This folder gets replaced when there are
upgrades.• Do not keep any of your programs in this folder.
![Page 14: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/14.jpg)
Exercise: Create a work shop folder
• Click New Folder button on Files Menu• Type in folder name• Click OK
![Page 15: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/15.jpg)
Exercise: Export a file• Check the box beside the folder you want to
export• Click the More drop down list• Choose Export
• Click download• Navigate to where you want to save the folder
• No need to export folder “data”
![Page 16: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/16.jpg)
Exercise: Manage Data Folder with Copy
• In Desktop/OTNSandbox• Right click folder ‘data’• Choose copy• Paste into same folder OTNsandbox• Rename copy to ‘data_cod_2008’• Delete everything from folder ‘data’
• If you open folder ‘data’ from Rstudio at this point it will be empty
![Page 17: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/17.jpg)
Exercise: Sample Data
• Open folder Desktop/OTNsandbox • Check that folder “data” is empty• If not, copy contents to another folder• Right click file:
SampleWorkShopData.zip• Choose Extract All
![Page 18: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/18.jpg)
Exercise: Sample Data
• Accept the file path • Check the box• Click Extract
![Page 19: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/19.jpg)
Exercise: Manage Data Folder
• In the Extract foldr open “data” folder • CNTL A – to highlight all• Right click • Choose Copy• Navigate to folder OTNsandbox/data• Paste
• Now if you open “data” folder in Rstudio all the workshop files should be there
![Page 20: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/20.jpg)
Documentation and Software Location • Introduction page with links
http://members.oceantrack.org/data/otn-tool-box
• Direct Location for most up to date Documentationhttp://members.oceantrack.org/toolbox/
![Page 21: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/21.jpg)
Folder Structure: Documentation
• There is extra stuff for geeks in the Appendix of the Install guide• Update Sandbox Tools Instructions would be used after initial install to add new
functions or fixes• Troubleshooting will be expanded as users report problems and we find solutions
![Page 22: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/22.jpg)
Exercise: File Conversion
• Open sandbox folder• Click on file_conversion_driver.r• File will open in upper left window of GUI
• Save file to WorkShop Scripts folder
![Page 23: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/23.jpg)
Exercise: File conversion
• Open data folder• Cut and paste the file
name into the script• Save the script
![Page 24: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/24.jpg)
Running R-scripts
• Highlight the lines you want to execute
• Click the run button
![Page 25: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/25.jpg)
File Conversion: NotePad++ Encoding
• Open file in NotePad++• Click Encoding on Menu Bar• Button indicates encoding• Click Convert to UTF-8 wo BOM• Save file
![Page 26: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/26.jpg)
False Filtering: Minimum Requirements• Column: unqdetecid must be present.• Must contain unique values.
• Column: catalognumber must be present.• This can be an animal id or a transmitter id.
• Column: datecollected must be present.• Must be format YYYY-MM-DD HH:MI:SS • or YYYY-MM-DDTHH:MI:SS• All digits must be present
• Column: station must be present.
![Page 27: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/27.jpg)
Exercise: filtering suspect detections
• Open sandbox folder• Click on filter_driver.r• Will open in upper left window• Save to WorkShop Scripts folder
![Page 28: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/28.jpg)
Exercise: Filtering Control Parameters
• Highlight this entire section and click the run button
![Page 29: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/29.jpg)
False Filtering: Set input values
• Open data folder• Highlight input detection file and copy• Paste into script window over detections.csv
![Page 30: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/30.jpg)
Exercise: Filtering Functions
loadDetections()• Input a detection file• Outputs a file of suspected detections• And an optional distance matrix
filterDetections()• Input a detection file and a file of suspect detections• Outputs a file of filtered detections • And an optional distance matrix
![Page 31: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/31.jpg)
Run the load step
• Paste the file name between the quotes• Highlight this section of code • Click the run button
![Page 32: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/32.jpg)
Output Messages: Load Step
![Page 33: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/33.jpg)
Data: Suspect Detections (transposed)
• Each row represents info about three consecutive detections of one animal• The column value for suspect_detection represents the unique id from the input file
![Page 34: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/34.jpg)
Run the Filter Step
• If you have your own file of suspect detections or have edited the one the tool created• This is where you override the input file• Otherwise the program will use the one created in the previous step
![Page 35: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/35.jpg)
Output Messages: Filter Step
• Messages will tell you: • What file of suspect detections was used• What the input detection file was• Record counts• Output file names
![Page 36: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/36.jpg)
Data: Distance Matrix
![Page 37: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/37.jpg)
Exercise: Distance Matrix Merge
• Open sandbox folder• Click on distance_matrix_merge_driver.r• Will open in upper left window• Save to WorkShop Scripts folder
![Page 38: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/38.jpg)
Exercise: Distance Matrix Merge
• File for distance_matrix_input was created in false filtering step• Highlight file sample_distance_matrix_override_values.csv in the data
folder and paste into distance_real_input expected value
![Page 39: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/39.jpg)
Exercise: Distance Matrix Merge
• Highlight entire script• Click Run
![Page 40: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/40.jpg)
Data: Distance Matrix
![Page 41: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/41.jpg)
Exercise: Interval Data
• Open sandbox folder• Click file interval_data_driver.r• Will open in upper left window• Save to WorkShop Scripts folder
![Page 42: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/42.jpg)
Exercise: Interval Data
• Grabbing filenames for subsequent steps• Find them in the output in the output Console. Bottom L.H.S.
![Page 43: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/43.jpg)
Exercise: Interval Data
• Grab the output file of detections from the last step of the filter step• Grab the output file from the distance matrix merge step• Paste values into the script• Save the script
![Page 44: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/44.jpg)
Exercise: Interval Data
• Execute the script
![Page 45: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/45.jpg)
Data: Compressed Detections
![Page 46: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/46.jpg)
Data: Interval Data
![Page 47: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/47.jpg)
Exercise: Cleanup
• Open sandbox folder• Click on file cleanup_driver.r• Will open in upper left window• Highlight entire script• Click Run
![Page 48: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/48.jpg)
Teach yourself to program
• Free open software• Extremely powerful• Standardized
• Python• Python(x,y): rival to MATLAB and Rstudio• PostgreSQL
![Page 49: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/49.jpg)
How? Coursera and Code Academy• Code Academy Python course:
http://www.codecademy.com/en/tracks/python
• Rice University: An Introduction to Interactive Programming in Python Next session Sep 15 https://www.coursera.org/course/interactivepython
• University of Michigan: Programming for Everybody Next Session Oct 6 https://www.coursera.org/course/pythonlearn
• Johns Hopkins: R Programming Part of the "Data Science" Specialization Next session Oct 6 https://www.coursera.org/course/rprog
![Page 50: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/50.jpg)
PostgreSQL: Online Tutorials
http://www.postgresqltutorial.com/
![Page 51: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/51.jpg)
RStudio vs R
• Prefer Rstudio https://www.rstudio.com/ide/download/• User friendly Interface
• Not standardized so use with caution• Null always TRUE• Unpredictable results• Unpredictable upgrades
• Help• Rseek: http://www.rseek.org/
![Page 52: HQ Workshop 2014](https://reader036.vdocument.in/reader036/viewer/2022062422/568135f2550346895d9d64db/html5/thumbnails/52.jpg)
Questions?
Wish list?• Cohort data• Separate files for animal detections on other lines• Station Group mapping function• If you can think it and describe it in English, we can
program it.