Wednesday, October 30, 2013

Kamehameha!

Dr. Remy and I recently discussed our long term goals for the Point Cloud Web Service project and our destination of the "jump rope" project. He wanted me to come up with a design for it and we discussed parameters such as the frequency of the rope's "jump", how it would actually move, and how we would determine a passing jump. It was definitely decided that the jump rope's swing should have a repeating pattern or period rather than swing at random intervals. This would aid the user with jumping over it successfully as without a prototype, we don't know how well the spatial audio will portray the moving rope in just one pass. Likewise we needed to choose between actually have the rope swing down below the feet and back up again or just move horizontally at a constant height. Would the user be able to tell the difference between the two methods if implemented with the Web Audio API through the browser? And would they be able to track the location of the rope more accurately using one or the other? Questions like these can only be answered with a prototype, but we can form hypotheses about them based off the Skeleton Tracking Audio program I've already implemented.

Dr. Remy also brought up one of my other ideas, basketball. We discussed how the user would locate the  ball, hold and move it around, and actually shoot it. My original idea was to have the ball be "invisible" and "silent" and for a shooting motion to initiate the sound of a shot, or a wooshing ball as it moves towards the basket and then a noise from the basket to indicate the accuracy and power of the shot versus how on target it was. We also pondered about allowing the user to control the ball before the shot and perhaps pass it around between each hand. This would aid with getting a feel for how your virtual ball behaved before you shot it as well as assist accuracy during the actual shot. This would be slightly more difficult though as it involves using the Kinect data to tell the difference between flinging the ball side to side versus a full shot straight forward on a perpendicular axis. We determined that perhaps a better method was to allow the user to hold  both hands together to "initialize" the ball and sound associated with it and then let them shoot at their leisure. If their hands move apart, the ball and sound disappear. You can't play around with the ball physics prior to the shot but you would still be able to see/hear how.

Talking about this design made me think of someone else's Kinect project that I had seen on youtube where they juggled several balls. Unfortunately I forgot to add it to my early blog post on example OpenKinect projects but as it turns out, it was done with OpenNI so that may have been why I forgot it. I've gone back and added it and here is a link to the juggling video. In the description is a link to his code as well. In his video it appears that he is using acceleration vectors from the motion of his hands to determine when/where to throw the balls. Otherwise when a ball collides with his hand it automatically sticks to it. Also, it appears that placing his hand out of sight and bringing it back in places a new ball in his hand. When balls move off screen they are removed and the color of a new ball is based on how many are on the screen at a given time. Interestingly the balls move in three dimensions as can be seen around the 50s mark in the video when he throws them at the camera and they grow largely in size (and at the beginning of the video when he demonstrates it). It's hard to determine if he must match the z-axis of his hand to the ball however in order to catch it as this would make catches much more difficult.

Out of curiosity, I followed the link in the description of his video to his "inspiration" which is this Kinect project that imitates a Kamehameha from Dragonball-Z. We were both intrigued by this other project as it is quite different from the more common "ball" or "object" tossing Kinect program. We also believe it might have some interesting applications to our work with point cloud web services and would be a better fit than jump rope, basketball, or juggling. The Kamehameha uses the same skeleton tracking API I was using (Microsoft Kinect for Windows SDK). The program highlights the persons outline with an "aura" and gives them the classic spiky hair from Dragonball-Z. When the user puts their hands together, a shiny orb is formed. Crouching and bringing the arms into a certain position makes the orb grow and when it is large enough, shifting your arms forward shoots the orb and results in a large beam of light (with some other fun side effects). Something to note is that it is very important the Kinect see your entire skeleton. The legs are very significant and without the entire view of your legs as you crouch, the orb won't grow very much and is hard to trigger.

Monday, October 28, 2013

More Papers

Dr. Remy took an interest in The Web Browser As Synthesizer And Interface and requested I do further research on it. The opening sentence from the abstract of the paper reads "our research examines the use and potential of native web technologies for musical expression". This is quite similar to our research on using the web browser as a mediator for other technologies. As such, it would be interesting to see what other papers this author has written, who has cited his paper, and what other papers share some of his sources.

The paper actually has three authors: Charlie Roberts, Graham Wakefield, and Matthew Wright. Wakefield is from a graduate school in Daejon, Republic of Korea while the other two are in the Media Arts and Technology program at University of California at Santa Barbara.

Charlie Roberts appears in another paper titled Enabling Multimodal Mobile Interfaces for Interactive
Musical Performance. The paper is about a mobile application called Control that provides a prototyping environment for defining multimodal interfaces that control real-time artistic and musical performances. He later "broadly" defines multimodal to be the combination of multiple signals from different modalities in order to generate a single coherent action. Control allows user to quickly create interfaces that use multiple modalities such as: speec recognition, computer vision, musical feature extraction, touchscreen widges, and inertial sensor data. The application uses JSON to create its interfaces and can be extended with JavaScript.

After discovering that Charlie Roberts did this second paper while at UCSB and with yet another person in the Media Arts program at USCB, I decided to find their publications page and see what else was going on in the program. A quick search on the page for "web" turns up the first paper I posted above as well as two others of potential interest. The first and most promising of which is www.lovelyweather.com: A Web-Based Interactive Audio-Visual Environment featuring two other UCSB students, Salman Bakht and Christopher Jette. The abstract explains that lovelyweather.com is a web-based hypermedia work created by the two authors and a third person, Alenjandro Casazi. Straight from the abstract "[the website] combines digital photographs with environmental audio recordings to form a virtual audio-visual space navigable by computer mouse. www.lovelyweather.com examines the relationsihp between the methdos of navigation used in web-based media and the process of exploration that an artist and audience experience in the creation, presentation, and interpretation of artwork". I found the abstract intriguing and visited www.lovelyweather.com myself. The project is quite interesting and essential starts you off in an environment, meaning the sounds and images displayed are related to the same physical environment. Certain "squares" or smaller images focus on a section of that environment and mousing over them appears to slowly change the audio to focus on that part of the environment. Likewise clicking them brings up a larger image of that section and re-focuses the audio. Spatial audio is also used in both the individual environment recordings themselves and when you are choosing a new environment to switch too.

The third paper from UCSB that mentions web browsers is Instructional Tools in Educational Measurement and Statistics (ITEMS) for School Personnel: Evaluation of Three Web-Based Training Modules. The paper focuses much more on the assessments themselves and the issues surrounding standardized testing rather than the medium, the browser, so I don't think this is very applicable to our research however it may still contain some relevant data or reflections so it is worth noting.

As for Charlie Robert's other publications from UCSB, he has quite a few, eleven to be exact. The latest two of which are those previously featured in this blog post. The third latest is Mobile Controls On-The-Fly:
An Abstraction for Distributed NIMEs which also makes mention of the application Control. The abstract notes that musical artists waste a lot of time on non-musical tasks such as network configuration, development environments for mobile devices, and OSC address spaces and environments. The authors' research is aimed at solving this problem by minimizing the code necessary to generate and read the output of interfaces on mobile devices.

Roberts also appears in two other papers of potential interest: Composition for Conductor and Audience:
New Uses for Mobile Devices in the Concert Hall and LuaAV: Extensibility and Heterogeneity for Audiovisual. The first is about an audience interaction piece called Composition for Conductor and Audience which leverages Control to allow audience members to become the "orchestra" by performing using different musical variables on the touchscreens of their mobile devices. The audience members actually downloaded the application just prior to the start of the concert and were given basic instructions on how to respond to certain gestures by the conductor in terms of an action on their device. I was able to find a video including some of the audio here. While it doesn't give a clear idea of what sound is being created by the conductor versus the audience, it shows visual examples of the application in use which is rather interesting. The second is a paper on the use of LuaAV, a runtime library that extends the Lua programing language, to support computational composition of temporal, sound, visual, spatial, and other elements. The abstract also states "In this paper we document how we have attempted to maintain several core principles of Lua itself... while providing the flexibility and temporal accuracy demanded by interactive audio-visual media".

Of the authors in the original paper that spawned this research tree, Graham Wakefield is featured in 24 papers from UCSB's Media Arts publications page. Titles of interest are "A Toolkit for Composing Immersive Audio-Visual Worlds of Agency and Autonomy", "Augmenting Computer Music with Just-In-Time Compilation", "The Allobrain: An Interactive, Stereographic, 3D Audio Immersive Virtual World", and "Experiencing Audio and Music in a Fully Immersive Environment".  Some of his others are shared with the Roberts papers previously discussed. Matthew Wright is also featured in 15 featuring "Sound Element Spatializer", a system for rendering and control of spatial audio. It operates on a cross-platform C++ application that can spatialize sound sources from other applications or live inputs in real-time.

I was unable to find any papers that cited the original paper that spurred this blogpost. It is relatively new (2013) and probably has not had time to engender others to further research.

As for the paper's sources, three other papers cite the source "C. Rogers. Web audio API. http://www.w3.org/TR/webaudio/" as seen here. The only paper of interest (other than the original) is yet another by Charlie Roberts. Gibber: Live coding audio in the browser is a paper that was mentioned on UCSB's Media Arts publications page but was not linked. Gibber is described as a javascript coding environment for web browsers that provides several synthesis options, audio effects, and sequencing objects to be controlled. It also enables networking of multiple users to perform simultaneously using a remotely controlled instance of Gibber.

Another source that I thought would have spawned other interesting papers is J. Young's J Using the web for live interactive music. The source itself is about a website that allows participants to navigate through a virtual interface and while manipulating elements receive a projection of their actions in the form of sounds in physical space. The live audio performance is streamed out to all participants and anyone could participate and hear the resulting music in real time.

Google scholar only finds two papers that cite this source, neither of which is C. Roberts original paper. I find "Public sound objects: a shared musical space on the web" by A. Barbosa the more interesting (the other also being in Spanish). Public Sound Objects is an on going project that allows users to collaborate as performers online. A public performance space is hosted along with the controls to create the audio. The mixture of all online collaborators is streamed back to the users as well any audience members and contextual data such as sound spatialization and metaphorical visual representations of current participants are added.

Finally, one of the sources of C. Roberts original paper that I found interesting was this one on Firefox and WebAudio. Though Mozilla doesn't support it yet, they are deprecating their Firefox specific Audio Data API and are in the process of adding support for the W3C standard Web Audio.

Monday, October 21, 2013

Web Requests - JSON vs. Text

After doing a fair amount of work involving web GET requests and transmission of data in different formats, I've taken an interest in the question of performance between the different types. Initially, I planned on comparing two ajax calls, one that used a response type of JSON and the other plain text. I was so used to using jQuery that I forgot that on a level lower than the common javascript library, response types might not exist. After doing some short research I discovered that the XMLHttpRequest object (the native javascript object used for web requests) has a responseType method through which you can set the data return type. This sounded promising but as it turns out, Chrome doesn't support the method yet. Since I am relying on the Audio API for spatial audio, Chrome is currently the only browser I'm concerned about so if there is no difference between a JSON and a plain text ajax call other than an extra call to JSON.parse(), the question of performance lacks any interest.

In fact, since the response type method doesn't work in Chrome, the definition of a "json" vs. "text" get request simply becomes a matter of how the data is parsed when it is received. But, in either case (response type or no response type) the question of JSON vs. plain text is bigger than the javascript. The data itself depends on the expected return type and JSON data may be much larger than a space delimited format, especially depending on how much metadata you wish that JSON to contain. But, there is another trade off on the clientside even after space delimited data is parsed, and that is how it is accessed and used. So aside from the XMLHttpRequest call itself, when comparing the performance of a "json" versus "text" style of GET requests, you also have to take into account the alternative method of parsing and the size of the extra metadata. For small transmissions the metadata is most likely negligible in terms of both network latency and parsing performance. But with large amounts of data with nested JSON would the performance difference be noticable? Would the JSON parsing also be slower than string splitting on spaces and simply storing in arrays? Would the performance results be different for small or large amounts of data based on any initial overhead of each method versus their respective scalability?

And this is just the client side. Whether the data being sent is expected to be JSON or plain text obviously affects how the server has to prepare the data. Does writing a string of JSON carry more overhead than simply concatenating a list of space delimited values? Similar to the client side questions, the server-side has several variables of its own.

Also, perhaps the value of the metadata, though it may not directly affect performance after it has been parsed, should add more value to the json method in terms of which is "best", not just in terms of performance. Having nicely named properties in an all-in-one object would certainly boost code readability and aid in the development process.

Monday, October 7, 2013

Skeleton Tracking Video + Commentary



Skeleton Tracking Requirements

So that someone might be able to run the program as well as modify it, here are links to the respective requirements.

Run-time Requirements:
SDK Requirements:
Also at the bottom of the HttpListener page in MSDN can be found the requirements for the C# web server class I am using. Of course the Kinect dependencies are the real limiting factors.

Skeleton Tracking - v1.0 Complete!

I've finally gotten the spatial audio + skeleton tracking model to work! After struggling with hosting the html webpage from webpy last week I realized I was overthinking the problem and could just host the html file on the same C# server by merely having C# read my html file and print the text to the screen. I got this working easily however two things were wrong, first, the mime-type was defaulting to text/plain and second the files I was including were erroring. The first thing wasn't really a problem, at least at the moment, because chrome still treated the text like html. The second problem was being caused because unlike with webpy, the C# server doesn't have a concept of a "static" folder so my requests for "static/jquery" etc. were just hitting the web server and receiving no response. I added responses for my jquery include but then realized I'd hit another issue with the cauldron file. Cauldron.wav needs to be sent as binary data and the vanilla C# web server I took from David's blog only provided string responses. At this point in time I also decided it'd be a good idea to go ahead and bundle all the files in Visual Studio so I added them as resources and could eliminate the need to open a file reader to them and have extra files to keep up with. VS Studio resources are quite easy to deal with and the index and jquery include were both stored as strings. The .wav file was stored as the datatype "UnmanagedMemoryStream". Using the .Read() method of a UMS I was able to write the byte data to a buffer. Then, I modified the C# server to pass the response method the actual response object which allowed me to do two things: 1) write byte data to the response 2) change the mime type. Thus I was able to respond to the request for the .wav file in binary format and also fixed the mimetype all 3 files. I came across this useful mapping of common extensions to mime types.

Additional changes I made to this first version of Skeleton Tracking were:
- Allowing for the starting/stopping of the server independent of the form (I had a button before but "stop" actually closed the server)
- Set the server to automatically start when the program does
- Renamed the major components of the program to "SkeletonTracking" instead of Microsoft's sample name since I have been directly modifying their skeleton tracking sample.

One final thing to note is that the webpage reading the kinect data is not actually using the filtered version of the data but the raw, noisy data. It didn't appear to have any affect on the spatial quality of the sound so I will leave it for now.

A video with audio and commentary is coming soon.

Noisy Skeleton (Video)

Here is a video of the type of noise I am experiencing. I was not wearing a jacket when testing and my arms were clearly visible. When my arms are far apart, the tracking is very smooth and there is almost no noise. But, when I put my hands close together in front of my stomach, you can see the wrist joints flicker dramatically. Even the elbow joints sometimes screw up. Likewise when I touch my shoulders.

Wednesday, October 2, 2013

Skeleton Tracking - HTML Frontend

Now that I've finished the C# server code and skeleton tracking, I have started the html front end to retrieve the kinect data and create spatial audio sound with it.

I was going to test natively on Windows however you can not make an ajax call using the file protocol and ajax is required to load the sound SO it was back to the VM to use webpy and a quick python server to host my html code. Even though I have no need (yet) for server-side python I at least had to have something hosting my html and sound file so that I could make an ajax request to it.

I've created a pretty basic html page that will simply make get requests to the C# server at the same interval as the previous ballbeam code did. The page has two buttons, one to stop the python service and one to start playing the audio. There are no visuals yet but the code to pan the audio based on the retrieved location is in place. After working through a few  bugs in the code, I have arrived at this error message:

Uncaught NetworkError: A network error occurred. X.cs.clemson.edu:8080/:64
XMLHttpRequest cannot load http://Y:8080/position. Origin http://X.cs.clemson.edu:8080 is not allowed by Access-Control-Allow-Origin. X.cs.clemson.edu:8080/:64


From this stack overflow question, it sounds like I need to add something to the header of my C# servers response to allow cross-domain requests from any (or at least X) domain. However, I can use "wget" to retrieve the C# server response data just fine from X. Perhaps wget uses a different outgoing port and that's why its able to get through? Either way I have two things to try: 1) I can still try setting the header data to allow all cross-domains as per stack overflow 2) I can look into why wget works and see if I can apply that too my ajax call.

Skeleton Tracking Server - Noisy Data

I forgot one more important piece to the server-side portion of the skeleton tracking server. The data provided by the Kinect is quite noisy and even in the SDK-provided visualization one can see the skeleton joints jumping around despite the tracked body being virtually motionless. Data this "noisy" would be lousy for audio panning as the audio would likely waver and vary at a high rate regardless of how the user moved their hand. To fix this a filter needs to be applied to the data (in this case simply rounding) so I will create to GET responses on the skeleton server, one for the unfiltered data and one for the rounded data.

C# Web Server + Skeleton Tracking

Once I became familiar with the skeleton tracking Kinect code, I began searching for a way to hook it up to a web service so I could access the skeleton data through http GET requests. The Kinect SDK's HTML5 examples were out of the question since they required Windows 8. Dr. Remy suggested the links below:

http://forums.asp.net/p/1597083/4053225.aspx
http://stackoverflow.com/questions/1140028/add-http-get-method-to-c-sharp-web-service
http://www.codeproject.com/Articles/863/Your-first-C-Web-Service
http://www.codeguru.com/csharp/csharp/cs_webservices/tutorials/article.php/c5477/NET-Web-Services-Tutorial.htm
http://support.microsoft.com/kb/308359

But unfortunately the only promising one's required a C# Visual Studio template called "ASP.NET Web Service". I tried searching all over for how to get this template. I started with Visual Studio 2012 and decided I might need VS Web 2012 to get this special template. After installing it I had access to several C#/ASP Templates however none of them were the "Web Service" template the tutorials kept mentioning. I then tried VS 2010 and Web 2010 with no luck. I did some searching on the "ASP.NET Web Service" template and found that it required .NET Framework 3.5 and that I had 4 and 4.5. I installed 3.5 but even then VS would not let me target the 3.5 framework when selecting from C# Web templates and I ultimately gave up on using Microsoft's templates.

Dr. Remy then recommended this C# server after some brief googling. I did some googling of my own and found a simpler one here. I threw it into my VS 2012 project and it worked like a charm, at least the sample did. Changing the server to output one of the skeleton joints instead of the sample string was very painful.

First, the web server class constructor required a static function to be used as the GET response handler. Making such a function static however cut of its access to the MainWindow class in my application which had the Kinect data. Then I added a method to the web server to allow setting a new response method and since it wasn't a constructor, the method didn't have to be static. Then I was able to declare a function in my main class (with the Kinect) that would return the Kinect data as a string which I hoped the web server would then respond with. However this failed due to a "cross-thread communication" exception. Then I created a static "DataPasser" class with global scope. In MainWindow I store the Kinect data in DataPasser and in the web sever response method I simply get the property from DataPasser. This worked however I didn't realize it at first since I had put the code that updates the DataPasser values in a function that is only called once the Kinect has detected a skeleton. Since I was running several trials of the program, I did not stand up and let the Kinect recognize my skeleton and thus no data (an empty web page) was returned to my requests.

After getting the data sorted out, I had to take the web server one step further and test it on another machine (up to this point I had done all testing on localhost). I found that from my laptop, which was on the same network as the desktop I was doing development on, I could not access the server. I did some googling and realized that the httplistener prefix "http://loclahost:8080/" only targets localhost and that "http://*:8080/index/" needed to be used to accept requests from any source. Upon trying this new prefix however VS was unable to run the program due to a failed call to the httplistener constructor. I did some more googling and found out this might be permissions related. Sure enough, restarting VS with admin rights fixed the issue and I was able to use the new prefix. However I was still unable to see the web server from my other laptop! Bringing down the firewall fixed this issue though so I added an exception to my firewall for the web server port and now it can be accessed from another machine on the network.

The next step in my development will now be to implement the HTML interface that will read the data and pan a sound based on the given position.