Wednesday, December 4, 2013

Web-Kamehameha with Beam

After receiving help on the geometry necessary to draw the beam from Dr. Remy, I was successfully able to write code to incorporate it into Web-Kamehameha. The below video shows the final product.


The solution to drawing the beam was to use the atan2 function (featured in almost every programming language) to get the angle of the center of the beam from the two given points. Then, using a factor of energy level and radians, I adjusted that angle to get the angle of the either side of the beam. From the newly computed angle, I determined which edge of the canvas the beam would hit (top, bottom, left, or right). Then, depending on the edge, I could determine the X or Y coordinate of the point to connect the base of the beam to. Given that X or Y coordinate, I could calculate it's distance from the opposite axis and use the tan() function and the basic formula tan = opposite / adjacent to find the length of the missing side, which would be directly equal to the missing coordinate. Below is my code.

function drawBeamSide(ballBase, angle) {
var adjacent = null;
var opposite = null;
var newX = ballBase[0];
var newY = ballBase[1];

// transpose the angle to the range [0, 2PI)
if (angle < 0) {
angle += 2 * Math.PI;
}

// determine which edge this beam should hit as it
// will affect the calculations
if (angle < 1 / 4 * Math.PI || angle > 7 / 4 * Math.PI) {
// Hits right edge
adjacent = ballBase[0] - 500; // distance from point to right edge
opposite = Math.tan(angle) * adjacent; // use trig to find Y coordinate
} else if (angle <= 3 / 4 * Math.PI) {
// Hits top edge
opposite = 0 - ballBase[1];
adjacent = opposite / Math.tan(angle);
} else if (angle <= 5 / 4 * Math.PI) {
// Hits left edge
adjacent = ballBase[0];
opposite = Math.tan(angle) * adjacent;
} else {
// Hits bottom edge
opposite = 500 - ballBase[1];
adjacent = opposite / Math.tan(angle);
}

newX -= adjacent;
newY += opposite;

drawLine(ballBase[0], ballBase[1], newX, newY);
}

The function's first argument is a two element array representing the X and Y coordinates of the ball's base (which is the same location for the beam's base) and then the angle (from the x-axis) of the beam to draw. Logic outside of this function already took care of changing the angle's and then calling this function to produce two sides to the beam.

The tricky thing to note is that the coordinate system for the HTML5 has a positive increasing Y value as you move down the image. The top of the image is Y = 0 while the bottom is Y = HEIGHT (or 500 in my case). The left is X = 0 and the right X = WIDTH.

Some interesting things to note about the program is that the depth appears to affect the accuracy or straightness of the beam. Tomoto's code, and what I've modified and put into my program, is simply drawinga vector along the line from the shoulder to the hand, however in the video below you will see that even when both are on the same Y level, the angle of the beam is quite off (too low). I noticed this effect increases with depth. It is not noticeable in the previous video since I crouched to increase the energy level. If I stand up straight and release a Kamehameha however you can see the problem.

In the below video I demonstrate this. I start off by showing the hand-distance threshold for ball generation. Then I release two crouched Kamehameha's for which the angle appears to be accurate. Then I release one standing and you can see that it shoots downward despite my arm being straight. I can even raise one hand to show the path of the beam, despite not having any energy, and you can see it is inaccurate.


Monday, December 2, 2013

Web-Kamehameha

The bare-bones web/claude based version of Kamehameha is almost complete. There is just one key piece of functionality missing, the energy beam! In the video below you will see that the user skeleton is being rendered and that when the hands come close enough together, an energy ball starts to grow. Things unnecessary to the core functionality, like dynamic lighting and the glowing auro, have been removed in favor of getting an example that simply proves the browser can be used to mediate the flow of data from the Kinect, to Kamehameha, to the browser.


When the energy ball is shot, by straightening the arms outward, currently all I have is a line showing the direction the beam should be going. The line is formed by the point at the center of the energy ball and what appears to be an arbitrarily chosen secondary point that forms a direction vector. I can't tell from Tomoto's code how the second point is calculated, however I haven't looked to far into it. His code also has no comments and is rather hard to follow. My idea for the beam is to have an angle (theta) that is a direct function of the energy level and represents how far a pair of lines is from the initial direction vector seen in the video. The lines can be of infinite length (to the edge of the screen) but they should each have a base at the ball center and one should be angle theta above the direction vector and the other theta below. I whipped up the below example in paint.
I thought I would be able to use basic geometry and trigonometry to find the X, Y, Z coordinates of the two lines that form the isosceles triangle to represent the beam, but it turned out to be harder than I thought. I basically need to use an angle and slope to find the slope of the lines for the beam, but then use that slope to calculate actual points. This problem can be thought of as 2-Dimensional since I am not using the Z dimension in my HTML5 rendering and it doesn't have any bearing on the visualization of the beam. But that didn't make it any easier as I couldn't figure out how to calculate the slope of a line theta degrees away from another line (sharing the circle center as a common base) nor how to take the slope of a line and given a pair of X, Y coordinates, find another pair of X, Y coordinates N units away from the first pair (where N is large enough to make the line go off screen). Or even more appropriately make the second pair of coordinates 0, Y where Y is a value such that 0, Y falls on that line, thus only drawing as much of the line as necessary.

Tuesday, November 26, 2013

Putting Force Data Online

To answer some questions from Dr. Remy's comment on the previous post, each Tactonic tile is 2'x2' with 24x24 sensels. Each sensel is one inch away from all adjacent sensels on the same panel. In the force array that the devices generate, the coordinates increase as follows:


The arrows indicate an increasing positive coordinate. It should be noted that when the array is printed, the results are not always consistent with this

I implemented a function to calculate coordinates of the center of pressure, as well, and it seems consistent with the pressure placed on the devices, though I have not had the chance to test this rigorously yet. The force data is also now printed to my localhost using Mongoose by Cesanta. The call to Mongoose prints a snapshot of the force array, and loading the page prints the force array stored at load time. Together, this data looks like so:


This is a record of me standing on the upper- and lower-middle-left of the master device (directions as indicated in above picture) with only the master device recording.

The program is set to run 3000 times, and the print function is only called when an HTTP Request is made. The print function prints out the run that the device is currently on, the force array for that run, and the center of pressure of that force array. The center of force is calculated when the print function is called.

Device 3A is currently not working, and when it is attached to the chain it causes a mirroring effect between 0A and 1A and also between 2A and 3A. The latter was discovered when I only put pressure on device 3A but pressure was recorded on 2A, which tells me that 3A can still determine force exerted upon it.

Mongoose Setup

The embeddable Mongoose server is distributed as mongoose.c and mongoose.h. All that is needed to use it is an #include "mongoose.h" in the file that uses the Mongoose API, and then include mongoose.c in the compile command. The server can then be started within a C or C++ program by calling the mg_start() function and passing it configuration parameters and an HTTP event handler callback function.


The mg_context struct holds the connection to the server once that connection is opened. The C-string array contains the configuration options needed to display the results of the Tactonic program. Currently, the port is set to 8080, the default for Mongoose, and the server root is being redirected to the location of the folder containing the program that calls the server (pardon the filepath names). C++ complains about the escaped spaces, but that is how spaces in the filepath must be represented on OSX.


Currently, the event callback just calls the printTouchRegistry() function, which has been altered to output a const char pointer instead of writing to file and to calculate the center of pressure when it is called.

The final piece is a call to mg_start().


This starts the server with the startup options and binds the event callback.

Currently, the server only posts a snapshot of the pressure data, as previously mentioned. If needed an asynchronous function can be made to update the pressure data in realtime. Mongoose works with Lua Server Pages, which could conceivably be used in conjunction with AJAX to update asynchronously.

Monday, November 25, 2013

Chrome vs. Firefox Performance

While working on the web version of Kamehameha, I've noticed a distinct performance difference between Chrome and Firefox. I don't know exactly what is causing the difference but the results can seen in this video. (Firefox on the left, Chrome on the right).


I am using the HTML5 canvas to render the skeleton however in the data for the joint locations come from GET requests fired every 50ms. It could be the HTML5 rendering engine, web request efficiency, or possibly something else that is accounting for the difference. To clarify, while the motion of my skeleton in firefox is smooth and consistent, in chrome it is very choppy and laggy. At times, chrome would run smooth too until it had been running for about 30s, then chrome would start to lag but firefox would still maintain its performance. Also note that the flashing mess of lines that appears in the lower left corner of each browser is the Kinect confusing my jacket, which is hanging on a chair in the background, for a skeleton.

Monday, November 18, 2013

Pressure Output

The output issue was solved today, and pressure output is now recorded from the Tactonic device. As of now, it writes the values to a text file whenever a nonzero force is recorded at every 20th call to the device. The output looks like this:

This is a record of me standing on two of the tiles.

The code for this operation doesn't do anything unusual: it grabs the recorded pressure values and outputs them to a file specified at the command line.


Whenever the device's touch callback is called, it calls the registerTouches() then check the result of that function and the number of times registerTouches() was called and runs printTouchRegistry() if writeFlag was true and the function has run some multiple of 20 times.

Sunday, November 17, 2013

Segfault Solved

The major roadblock I encountered writing code to extract raw data from the Tactonic device was a segmentation fault in the startup phase. When the code attempted to assign a device to the custom visualizer class, execution would immediately stop due to an EXC_BAD_ACCESS error being thrown (see debug output below).

The instance of the custom viewer was originally a viewer pointer, but after changing it to a non-pointer instance the EXC_BAD_ACCESS error was no longer thrown, and the device paired with the viewer normally. I managed to output the force matrix generated by the Tactonic device, and it now gets written to a text file supplied as an argument. I'm working on a mainloop that will output the force matrix at every x calls to change it, but for now the program outputs this:
When running through the command line, a record of changing pressures was left, showing that the device is still recognizing and outputting nonzero pressures.

Once the mainloop is finished, we'll have a better idea of what needs to be done to get meaningful data from the device.

Tuesday, November 12, 2013

Video Reaction

The demo video of Tactonic Technologies' Tactonic Sensors was enlightening despite the flaws in the presentation itself. Their product is a malleable pressure sensor that has nodes spaced at 1/2", yet they only created a visualizer to see pressure patterns on the sensor. I agree that this has its uses--they mentioned medical and commercial applications, and one that comes to mind for me is determining the shape of and weight distribution across a person's foot--but they did not provide a set of functions for extracting raw data from the pressure sensor. The presentation also did not talk about the technical side of the sensors much, but this video seems to have been more of a sales pitch than a tech talk so that's understandable.

The first complaint above (not providing a means of extracting raw data) is the subject of my first investigation. I will be attempting to create a library that allows users to get raw data instead of images when they use the pressure sensors. The majority of the work will be in interpreting the provided source classes to see how they gather data.

Monday, November 11, 2013

Compiling Kamehameha

Before I can make changes to Kamehaha, I need to first get a working development environment for it. I was hopefully that since the google code project had a directory called "build" with a microsoft visual studio solution file (.sln) in it that it would work with visual studio. The file is from an older version than MSVS C++ 2010 which I used to convert and open it. Upon trying to compile the program however, I received a serious of errors due to the missing header file "XnCppWrapper.h". A quick google showed me this was a header from the OpenNI library. This is puzzling because the windows version of Kamehaha was supposed to be using the Microsoft SDK according to the project page. However it looks like there is an OpenNI port as well, not just for linux but even for windows. It also looks like the build2010 folder in the source files has another microsoft visual studio solution which is perhaps the Kinect SDK one. Since I will be modifying the kamehameha program to use the web service for kinect data, it doesn't really matter which library the solution is using for now since I will eventually eliminate that dependency. But, to try and just get the solution to compile, I went ahead and installed the OpenNI sdk. This however didn't fix the issue and upon searching the installation directory for the missing header file I saw that it had not been installed. My google of the wrapper had turned up this github which is what initially told me it was part of OpenNI.

At this point, I looked back to the google code project for kinect kamehameha to see if I missed anything. I did realize that I had checked out just the trunk from the svn repository and reliazed there were 3 separate branches with what I thought might be different code. They were: dev-1.0a, openni2-beta, and port-linux. Logically dev-1.0a seemd the only possibility for the Microsoft Kinect SDK but alas it too required the XnCppWrapper.h header file. Also upon commenting out the include a host of errors appeared as clearly the code was using the OpenNI library. So the mystery remains, where is the Microsoft Kinect SDK version of the code that the author so clearly states he is using? There was a second solution (.sln) file in the "build2010" folder and even though it didn't have any source files, which meant it was probably using the same files as "build" which were pulled from the "src" folder which are using OpenNI, I decided to try it. It actually compiled fine but when it came time to run the executable, windows threw an error saying the application failed to start due to the application config being incorrect. Googling this problem was useless as a whole host of different things can cause it.

Finally I decided to try compiling the source files with MinGW. This however yields an error about the missing header file "crtdbg.h" which is apparently a visual studio file necessary for debugging memory leaks. This include is in nearly every source file so it would not be trivial to just comment it completely out and see if it works.

So, with not many options left, I went back to the github which had the missing header file needed by openNI and downloaded the entire repository. I then linked the visual studio project to the downloaded library and it resolved that particular missing header error. However, I was still getting an error due to miss-matching path and output file names between the linker and some other configuration file in visual studio. I renamed the linkers output file name to match the other value (not sure where it was coming from as VS just said "targetname:") and now have one error left, LINK: cannot open file 'OpenNI.lib'.

I did some googling and found this post. At the bottom, someone mentions checking the path variable. I didn't have that exact path variable set but nothing in the error message told me to what the actual name should be so instead I looked at the linker's dependencies and found the "OpenNI.lib" dependency. It was relative so I went to the OpenNI install directory from my installation earlier and found OpenNI2.lib. Due to a newer version of OpenNI, it looks like it was not picking up the library. I then changed the dependency to OpenNI2.lib and added its directory to the "Additional Library Directories" setting under Linker > General in VS. Unfortunately it looks like OpenNI2 had lots of changes as there were multiple compiler errors to unresolved symbols which I'm guessing were removed or changed in OpenNI2.

So, I proceeded to install the latest version of OpenNI (1). I remembered he had a branch called OpenNI2-beta but since it was "beta" decided not to try my luck and have to start all over trying to get the VS solution to compile. So, with openNI (v1.5.7.8) installed alongside openNI2, I switched the solution back to using the old lib. At this point I also realized the github project I had downloaded was also v1.5.7.8 of openNI and that could have been part of the reason the previous compile didn't work.

At any rate, finally, the compile successfully completed... only to fail when VS tried to run the executable and producing the same error as running the executable produced from the solution in the "build2010" folder.

Then I began googling the error message again. I found this post which doesn't seem to relate much to my issue however it mentioned dependency's and a dependency walking program. A quick google and I found this program which when used on the executable I was producing, discovered the program could not find 3 DLLs:

OPENNI.DLL
GPSVC.DLL
IESHIMS.DLL

I've been trying to resolve the openNI issue first but haven't turned up much. I lost the link but an article I read on the issue mentioned that your should be including the appropriate ".lib" file in the visual studio project and that the .lib file is responsible for linking the executable to the dll. Unfortunately I already am including the correct .lib file so it sounds like perhaps its an openNI issue... Also, I was able to find the openni.dll under the openNI installation directory in the "bin" folder. I tried adding that folder to my path but it made no difference on the executable and dependency walker still showed it couldn't find it.

Then I found a way to get dependency walker to give me the path that the program was using to look for the DLLs and found that the path for all 3 of the missing DLLs was simply the debug folder i.e. the output folder for the executable.

I did some research and found this article which leads me to believe it is a 32 vs. 64 bit issue. I installed the openNI 32-bit library so perhaps that was part of it. I tried manually finding and copying the missing dlls into the debug folder but then dependency walker said there were "conflicting CPU types".

Unfortunately, installing the 64-bit version of openNI didn't work either as I got similar unresolved external symbol errors I received when I tried openNI2.

To be thorough, I went back to the solution from the build2010 file and ran dependency walker on the executable it produced. This time the openni.dll was not missing but the other two previously mentioned were. It is also worth noting that the gpsvc.dll and ieshims.dlls are related to group policy and internet explorer respectively. Two dlls that have little to do with this project, so I think. Also dependency walker includes an extra icon next to each of these entires that openni.dll didn't have when it was missing from the other solution. It looks like they might be delay load dependencies. The second question on their website's faq here seems to indicate that those types of missing dependencies are harmless so long as the calling dll handles it properly. In all likely hood I'm guessing those two missing dlls have nothing to do with the executable's issue to run.

So, back to the drawing board. This time, I noticed that at the end of the error message from running the executable, windows says "check the Application error log for more information". Well, as it turns out, the application log does have some useful information. In particular, it was logging the following:

Activation context generation failed for "C:\Users\DARA\Desktop\kamehameha\dev-1.0a\build2010\MSSDK_Debug\kinect-kamehameha.exe". Dependent Assembly Microsoft.VC90.DebugCRT,processorArchitecture="x86",publicKeyToken="XXXXXXX",type="win32",version="9.0.21022.8" could not be found. Please use sxstrace.exe for detailed diagnosis.

I started googling for sxstrace.exe thinking it would give me more details and found this article. A key thing to note here is that in the application log, the source of the error was "SideBySide". Also, that article I linked to led me to understand that Microsoft.VC90.DebugCRT was causing the problem. Googling that led me to this post and ultimately, the final answer. And boy, is it anti-climatic. The first reply states this:

It looks like you have a dependency on a debug library (Microsoft.VC90.DebugCRT)
Debug libraries are not provided in any redistributable pack (because they are not redistributable).
This probably means you are trying to run a debug build on a machine that does not have Visual Studio installed. The solution to the problem is to  build and distribute a release build, not a debug build.
I had noticed that the project was set to Debug but I thought nothing of it. I had no idea that a "debug" version could cause such issues. Simply changing the drop down in visual studio (there is no name or way to describe said dropdown other than it is in the top toolbar...) to _Release fixed everything. The program compiled and ran flawlessly. Here I will also note that they are two versions of each _debug _release pair respectively, MSSDK_ and OpenNI_. The "build" folder only had debug and release and since I was required to do all of the crazy include stuff for OpenNI, I guess that was the only library it supported. The build2010 folder however worked find without it and I now see that is because it lets you choose between which library to use.
So this long journey through dependencies and libraries ended up being a single value fix. From debug to release. Quite frustrating. Looking back there is not much I could've done to get to the solution quicker either. Perhaps noticing that the application logs had more information on the problem would've saved me the time of trying to fix the "build" folders dependencies.

Wednesday, November 6, 2013

Songs of Diridum

Songs of Diridum is a cool 3D exploration/experience demo for an upcoming browser based fantasy rpg game. The demo currently only allows you to move around in a small village area and listen to the playing jazz band however they have their sites set big for the full release. The most interesting part about the demo is that it requires no plug-ins, additional software, or dependencies other than HTML5 and WebAudio support, which is a standard for most modern browsers. The demo features quirky, blocky structures and people, most of which are inanimate. But, on stage is a band swaying to the music as they play a soft jazz tune. As previously mentioned, the WebAudio API is leveraged in the gaming platform used to create this demo (goo) and you can walk up to the band and hear the difference in volume and spatial distance. You can also make your character jump and after hopping onto the stage you can walk around and really take-in the full effects of the WebAudio APIs spatial audio. Getting closer to one instrument distinctly brings it out and even the way you are facing has a large impact on the sound. As discussed in the link at the beginning of this post, the players each have their own audio cone to make their sounds more realistic based on where your character is. Both the graphics and spatial audio are packaged together in a web-based gaming platform called goo. Goo itself is built on top of the WebGL library but adds things like spatial audio so that you can configure it all from one API. It will certainly be interesting to see what future browser games come out thanks to goo. It may also be very useful to our purposes for using the browser as a unified, cross-platform medium for communication. In particular, since kamehameha will require visuals of some kind, goo might be a good candidate for rendering them while maintaining our dependency-free point cloud system.

Wednesday, October 30, 2013

Kamehameha!

Dr. Remy and I recently discussed our long term goals for the Point Cloud Web Service project and our destination of the "jump rope" project. He wanted me to come up with a design for it and we discussed parameters such as the frequency of the rope's "jump", how it would actually move, and how we would determine a passing jump. It was definitely decided that the jump rope's swing should have a repeating pattern or period rather than swing at random intervals. This would aid the user with jumping over it successfully as without a prototype, we don't know how well the spatial audio will portray the moving rope in just one pass. Likewise we needed to choose between actually have the rope swing down below the feet and back up again or just move horizontally at a constant height. Would the user be able to tell the difference between the two methods if implemented with the Web Audio API through the browser? And would they be able to track the location of the rope more accurately using one or the other? Questions like these can only be answered with a prototype, but we can form hypotheses about them based off the Skeleton Tracking Audio program I've already implemented.

Dr. Remy also brought up one of my other ideas, basketball. We discussed how the user would locate the  ball, hold and move it around, and actually shoot it. My original idea was to have the ball be "invisible" and "silent" and for a shooting motion to initiate the sound of a shot, or a wooshing ball as it moves towards the basket and then a noise from the basket to indicate the accuracy and power of the shot versus how on target it was. We also pondered about allowing the user to control the ball before the shot and perhaps pass it around between each hand. This would aid with getting a feel for how your virtual ball behaved before you shot it as well as assist accuracy during the actual shot. This would be slightly more difficult though as it involves using the Kinect data to tell the difference between flinging the ball side to side versus a full shot straight forward on a perpendicular axis. We determined that perhaps a better method was to allow the user to hold  both hands together to "initialize" the ball and sound associated with it and then let them shoot at their leisure. If their hands move apart, the ball and sound disappear. You can't play around with the ball physics prior to the shot but you would still be able to see/hear how.

Talking about this design made me think of someone else's Kinect project that I had seen on youtube where they juggled several balls. Unfortunately I forgot to add it to my early blog post on example OpenKinect projects but as it turns out, it was done with OpenNI so that may have been why I forgot it. I've gone back and added it and here is a link to the juggling video. In the description is a link to his code as well. In his video it appears that he is using acceleration vectors from the motion of his hands to determine when/where to throw the balls. Otherwise when a ball collides with his hand it automatically sticks to it. Also, it appears that placing his hand out of sight and bringing it back in places a new ball in his hand. When balls move off screen they are removed and the color of a new ball is based on how many are on the screen at a given time. Interestingly the balls move in three dimensions as can be seen around the 50s mark in the video when he throws them at the camera and they grow largely in size (and at the beginning of the video when he demonstrates it). It's hard to determine if he must match the z-axis of his hand to the ball however in order to catch it as this would make catches much more difficult.

Out of curiosity, I followed the link in the description of his video to his "inspiration" which is this Kinect project that imitates a Kamehameha from Dragonball-Z. We were both intrigued by this other project as it is quite different from the more common "ball" or "object" tossing Kinect program. We also believe it might have some interesting applications to our work with point cloud web services and would be a better fit than jump rope, basketball, or juggling. The Kamehameha uses the same skeleton tracking API I was using (Microsoft Kinect for Windows SDK). The program highlights the persons outline with an "aura" and gives them the classic spiky hair from Dragonball-Z. When the user puts their hands together, a shiny orb is formed. Crouching and bringing the arms into a certain position makes the orb grow and when it is large enough, shifting your arms forward shoots the orb and results in a large beam of light (with some other fun side effects). Something to note is that it is very important the Kinect see your entire skeleton. The legs are very significant and without the entire view of your legs as you crouch, the orb won't grow very much and is hard to trigger.

Monday, October 28, 2013

More Papers

Dr. Remy took an interest in The Web Browser As Synthesizer And Interface and requested I do further research on it. The opening sentence from the abstract of the paper reads "our research examines the use and potential of native web technologies for musical expression". This is quite similar to our research on using the web browser as a mediator for other technologies. As such, it would be interesting to see what other papers this author has written, who has cited his paper, and what other papers share some of his sources.

The paper actually has three authors: Charlie Roberts, Graham Wakefield, and Matthew Wright. Wakefield is from a graduate school in Daejon, Republic of Korea while the other two are in the Media Arts and Technology program at University of California at Santa Barbara.

Charlie Roberts appears in another paper titled Enabling Multimodal Mobile Interfaces for Interactive
Musical Performance. The paper is about a mobile application called Control that provides a prototyping environment for defining multimodal interfaces that control real-time artistic and musical performances. He later "broadly" defines multimodal to be the combination of multiple signals from different modalities in order to generate a single coherent action. Control allows user to quickly create interfaces that use multiple modalities such as: speec recognition, computer vision, musical feature extraction, touchscreen widges, and inertial sensor data. The application uses JSON to create its interfaces and can be extended with JavaScript.

After discovering that Charlie Roberts did this second paper while at UCSB and with yet another person in the Media Arts program at USCB, I decided to find their publications page and see what else was going on in the program. A quick search on the page for "web" turns up the first paper I posted above as well as two others of potential interest. The first and most promising of which is www.lovelyweather.com: A Web-Based Interactive Audio-Visual Environment featuring two other UCSB students, Salman Bakht and Christopher Jette. The abstract explains that lovelyweather.com is a web-based hypermedia work created by the two authors and a third person, Alenjandro Casazi. Straight from the abstract "[the website] combines digital photographs with environmental audio recordings to form a virtual audio-visual space navigable by computer mouse. www.lovelyweather.com examines the relationsihp between the methdos of navigation used in web-based media and the process of exploration that an artist and audience experience in the creation, presentation, and interpretation of artwork". I found the abstract intriguing and visited www.lovelyweather.com myself. The project is quite interesting and essential starts you off in an environment, meaning the sounds and images displayed are related to the same physical environment. Certain "squares" or smaller images focus on a section of that environment and mousing over them appears to slowly change the audio to focus on that part of the environment. Likewise clicking them brings up a larger image of that section and re-focuses the audio. Spatial audio is also used in both the individual environment recordings themselves and when you are choosing a new environment to switch too.

The third paper from UCSB that mentions web browsers is Instructional Tools in Educational Measurement and Statistics (ITEMS) for School Personnel: Evaluation of Three Web-Based Training Modules. The paper focuses much more on the assessments themselves and the issues surrounding standardized testing rather than the medium, the browser, so I don't think this is very applicable to our research however it may still contain some relevant data or reflections so it is worth noting.

As for Charlie Robert's other publications from UCSB, he has quite a few, eleven to be exact. The latest two of which are those previously featured in this blog post. The third latest is Mobile Controls On-The-Fly:
An Abstraction for Distributed NIMEs which also makes mention of the application Control. The abstract notes that musical artists waste a lot of time on non-musical tasks such as network configuration, development environments for mobile devices, and OSC address spaces and environments. The authors' research is aimed at solving this problem by minimizing the code necessary to generate and read the output of interfaces on mobile devices.

Roberts also appears in two other papers of potential interest: Composition for Conductor and Audience:
New Uses for Mobile Devices in the Concert Hall and LuaAV: Extensibility and Heterogeneity for Audiovisual. The first is about an audience interaction piece called Composition for Conductor and Audience which leverages Control to allow audience members to become the "orchestra" by performing using different musical variables on the touchscreens of their mobile devices. The audience members actually downloaded the application just prior to the start of the concert and were given basic instructions on how to respond to certain gestures by the conductor in terms of an action on their device. I was able to find a video including some of the audio here. While it doesn't give a clear idea of what sound is being created by the conductor versus the audience, it shows visual examples of the application in use which is rather interesting. The second is a paper on the use of LuaAV, a runtime library that extends the Lua programing language, to support computational composition of temporal, sound, visual, spatial, and other elements. The abstract also states "In this paper we document how we have attempted to maintain several core principles of Lua itself... while providing the flexibility and temporal accuracy demanded by interactive audio-visual media".

Of the authors in the original paper that spawned this research tree, Graham Wakefield is featured in 24 papers from UCSB's Media Arts publications page. Titles of interest are "A Toolkit for Composing Immersive Audio-Visual Worlds of Agency and Autonomy", "Augmenting Computer Music with Just-In-Time Compilation", "The Allobrain: An Interactive, Stereographic, 3D Audio Immersive Virtual World", and "Experiencing Audio and Music in a Fully Immersive Environment".  Some of his others are shared with the Roberts papers previously discussed. Matthew Wright is also featured in 15 featuring "Sound Element Spatializer", a system for rendering and control of spatial audio. It operates on a cross-platform C++ application that can spatialize sound sources from other applications or live inputs in real-time.

I was unable to find any papers that cited the original paper that spurred this blogpost. It is relatively new (2013) and probably has not had time to engender others to further research.

As for the paper's sources, three other papers cite the source "C. Rogers. Web audio API. http://www.w3.org/TR/webaudio/" as seen here. The only paper of interest (other than the original) is yet another by Charlie Roberts. Gibber: Live coding audio in the browser is a paper that was mentioned on UCSB's Media Arts publications page but was not linked. Gibber is described as a javascript coding environment for web browsers that provides several synthesis options, audio effects, and sequencing objects to be controlled. It also enables networking of multiple users to perform simultaneously using a remotely controlled instance of Gibber.

Another source that I thought would have spawned other interesting papers is J. Young's J Using the web for live interactive music. The source itself is about a website that allows participants to navigate through a virtual interface and while manipulating elements receive a projection of their actions in the form of sounds in physical space. The live audio performance is streamed out to all participants and anyone could participate and hear the resulting music in real time.

Google scholar only finds two papers that cite this source, neither of which is C. Roberts original paper. I find "Public sound objects: a shared musical space on the web" by A. Barbosa the more interesting (the other also being in Spanish). Public Sound Objects is an on going project that allows users to collaborate as performers online. A public performance space is hosted along with the controls to create the audio. The mixture of all online collaborators is streamed back to the users as well any audience members and contextual data such as sound spatialization and metaphorical visual representations of current participants are added.

Finally, one of the sources of C. Roberts original paper that I found interesting was this one on Firefox and WebAudio. Though Mozilla doesn't support it yet, they are deprecating their Firefox specific Audio Data API and are in the process of adding support for the W3C standard Web Audio.

Monday, October 21, 2013

Web Requests - JSON vs. Text

After doing a fair amount of work involving web GET requests and transmission of data in different formats, I've taken an interest in the question of performance between the different types. Initially, I planned on comparing two ajax calls, one that used a response type of JSON and the other plain text. I was so used to using jQuery that I forgot that on a level lower than the common javascript library, response types might not exist. After doing some short research I discovered that the XMLHttpRequest object (the native javascript object used for web requests) has a responseType method through which you can set the data return type. This sounded promising but as it turns out, Chrome doesn't support the method yet. Since I am relying on the Audio API for spatial audio, Chrome is currently the only browser I'm concerned about so if there is no difference between a JSON and a plain text ajax call other than an extra call to JSON.parse(), the question of performance lacks any interest.

In fact, since the response type method doesn't work in Chrome, the definition of a "json" vs. "text" get request simply becomes a matter of how the data is parsed when it is received. But, in either case (response type or no response type) the question of JSON vs. plain text is bigger than the javascript. The data itself depends on the expected return type and JSON data may be much larger than a space delimited format, especially depending on how much metadata you wish that JSON to contain. But, there is another trade off on the clientside even after space delimited data is parsed, and that is how it is accessed and used. So aside from the XMLHttpRequest call itself, when comparing the performance of a "json" versus "text" style of GET requests, you also have to take into account the alternative method of parsing and the size of the extra metadata. For small transmissions the metadata is most likely negligible in terms of both network latency and parsing performance. But with large amounts of data with nested JSON would the performance difference be noticable? Would the JSON parsing also be slower than string splitting on spaces and simply storing in arrays? Would the performance results be different for small or large amounts of data based on any initial overhead of each method versus their respective scalability?

And this is just the client side. Whether the data being sent is expected to be JSON or plain text obviously affects how the server has to prepare the data. Does writing a string of JSON carry more overhead than simply concatenating a list of space delimited values? Similar to the client side questions, the server-side has several variables of its own.

Also, perhaps the value of the metadata, though it may not directly affect performance after it has been parsed, should add more value to the json method in terms of which is "best", not just in terms of performance. Having nicely named properties in an all-in-one object would certainly boost code readability and aid in the development process.

Monday, October 7, 2013

Skeleton Tracking Video + Commentary



Skeleton Tracking Requirements

So that someone might be able to run the program as well as modify it, here are links to the respective requirements.

Run-time Requirements:
SDK Requirements:
Also at the bottom of the HttpListener page in MSDN can be found the requirements for the C# web server class I am using. Of course the Kinect dependencies are the real limiting factors.

Skeleton Tracking - v1.0 Complete!

I've finally gotten the spatial audio + skeleton tracking model to work! After struggling with hosting the html webpage from webpy last week I realized I was overthinking the problem and could just host the html file on the same C# server by merely having C# read my html file and print the text to the screen. I got this working easily however two things were wrong, first, the mime-type was defaulting to text/plain and second the files I was including were erroring. The first thing wasn't really a problem, at least at the moment, because chrome still treated the text like html. The second problem was being caused because unlike with webpy, the C# server doesn't have a concept of a "static" folder so my requests for "static/jquery" etc. were just hitting the web server and receiving no response. I added responses for my jquery include but then realized I'd hit another issue with the cauldron file. Cauldron.wav needs to be sent as binary data and the vanilla C# web server I took from David's blog only provided string responses. At this point in time I also decided it'd be a good idea to go ahead and bundle all the files in Visual Studio so I added them as resources and could eliminate the need to open a file reader to them and have extra files to keep up with. VS Studio resources are quite easy to deal with and the index and jquery include were both stored as strings. The .wav file was stored as the datatype "UnmanagedMemoryStream". Using the .Read() method of a UMS I was able to write the byte data to a buffer. Then, I modified the C# server to pass the response method the actual response object which allowed me to do two things: 1) write byte data to the response 2) change the mime type. Thus I was able to respond to the request for the .wav file in binary format and also fixed the mimetype all 3 files. I came across this useful mapping of common extensions to mime types.

Additional changes I made to this first version of Skeleton Tracking were:
- Allowing for the starting/stopping of the server independent of the form (I had a button before but "stop" actually closed the server)
- Set the server to automatically start when the program does
- Renamed the major components of the program to "SkeletonTracking" instead of Microsoft's sample name since I have been directly modifying their skeleton tracking sample.

One final thing to note is that the webpage reading the kinect data is not actually using the filtered version of the data but the raw, noisy data. It didn't appear to have any affect on the spatial quality of the sound so I will leave it for now.

A video with audio and commentary is coming soon.

Noisy Skeleton (Video)

Here is a video of the type of noise I am experiencing. I was not wearing a jacket when testing and my arms were clearly visible. When my arms are far apart, the tracking is very smooth and there is almost no noise. But, when I put my hands close together in front of my stomach, you can see the wrist joints flicker dramatically. Even the elbow joints sometimes screw up. Likewise when I touch my shoulders.

Wednesday, October 2, 2013

Skeleton Tracking - HTML Frontend

Now that I've finished the C# server code and skeleton tracking, I have started the html front end to retrieve the kinect data and create spatial audio sound with it.

I was going to test natively on Windows however you can not make an ajax call using the file protocol and ajax is required to load the sound SO it was back to the VM to use webpy and a quick python server to host my html code. Even though I have no need (yet) for server-side python I at least had to have something hosting my html and sound file so that I could make an ajax request to it.

I've created a pretty basic html page that will simply make get requests to the C# server at the same interval as the previous ballbeam code did. The page has two buttons, one to stop the python service and one to start playing the audio. There are no visuals yet but the code to pan the audio based on the retrieved location is in place. After working through a few  bugs in the code, I have arrived at this error message:

Uncaught NetworkError: A network error occurred. X.cs.clemson.edu:8080/:64
XMLHttpRequest cannot load http://Y:8080/position. Origin http://X.cs.clemson.edu:8080 is not allowed by Access-Control-Allow-Origin. X.cs.clemson.edu:8080/:64


From this stack overflow question, it sounds like I need to add something to the header of my C# servers response to allow cross-domain requests from any (or at least X) domain. However, I can use "wget" to retrieve the C# server response data just fine from X. Perhaps wget uses a different outgoing port and that's why its able to get through? Either way I have two things to try: 1) I can still try setting the header data to allow all cross-domains as per stack overflow 2) I can look into why wget works and see if I can apply that too my ajax call.

Skeleton Tracking Server - Noisy Data

I forgot one more important piece to the server-side portion of the skeleton tracking server. The data provided by the Kinect is quite noisy and even in the SDK-provided visualization one can see the skeleton joints jumping around despite the tracked body being virtually motionless. Data this "noisy" would be lousy for audio panning as the audio would likely waver and vary at a high rate regardless of how the user moved their hand. To fix this a filter needs to be applied to the data (in this case simply rounding) so I will create to GET responses on the skeleton server, one for the unfiltered data and one for the rounded data.

C# Web Server + Skeleton Tracking

Once I became familiar with the skeleton tracking Kinect code, I began searching for a way to hook it up to a web service so I could access the skeleton data through http GET requests. The Kinect SDK's HTML5 examples were out of the question since they required Windows 8. Dr. Remy suggested the links below:

http://forums.asp.net/p/1597083/4053225.aspx
http://stackoverflow.com/questions/1140028/add-http-get-method-to-c-sharp-web-service
http://www.codeproject.com/Articles/863/Your-first-C-Web-Service
http://www.codeguru.com/csharp/csharp/cs_webservices/tutorials/article.php/c5477/NET-Web-Services-Tutorial.htm
http://support.microsoft.com/kb/308359

But unfortunately the only promising one's required a C# Visual Studio template called "ASP.NET Web Service". I tried searching all over for how to get this template. I started with Visual Studio 2012 and decided I might need VS Web 2012 to get this special template. After installing it I had access to several C#/ASP Templates however none of them were the "Web Service" template the tutorials kept mentioning. I then tried VS 2010 and Web 2010 with no luck. I did some searching on the "ASP.NET Web Service" template and found that it required .NET Framework 3.5 and that I had 4 and 4.5. I installed 3.5 but even then VS would not let me target the 3.5 framework when selecting from C# Web templates and I ultimately gave up on using Microsoft's templates.

Dr. Remy then recommended this C# server after some brief googling. I did some googling of my own and found a simpler one here. I threw it into my VS 2012 project and it worked like a charm, at least the sample did. Changing the server to output one of the skeleton joints instead of the sample string was very painful.

First, the web server class constructor required a static function to be used as the GET response handler. Making such a function static however cut of its access to the MainWindow class in my application which had the Kinect data. Then I added a method to the web server to allow setting a new response method and since it wasn't a constructor, the method didn't have to be static. Then I was able to declare a function in my main class (with the Kinect) that would return the Kinect data as a string which I hoped the web server would then respond with. However this failed due to a "cross-thread communication" exception. Then I created a static "DataPasser" class with global scope. In MainWindow I store the Kinect data in DataPasser and in the web sever response method I simply get the property from DataPasser. This worked however I didn't realize it at first since I had put the code that updates the DataPasser values in a function that is only called once the Kinect has detected a skeleton. Since I was running several trials of the program, I did not stand up and let the Kinect recognize my skeleton and thus no data (an empty web page) was returned to my requests.

After getting the data sorted out, I had to take the web server one step further and test it on another machine (up to this point I had done all testing on localhost). I found that from my laptop, which was on the same network as the desktop I was doing development on, I could not access the server. I did some googling and realized that the httplistener prefix "http://loclahost:8080/" only targets localhost and that "http://*:8080/index/" needed to be used to accept requests from any source. Upon trying this new prefix however VS was unable to run the program due to a failed call to the httplistener constructor. I did some more googling and found out this might be permissions related. Sure enough, restarting VS with admin rights fixed the issue and I was able to use the new prefix. However I was still unable to see the web server from my other laptop! Bringing down the firewall fixed this issue though so I added an exception to my firewall for the web server port and now it can be accessed from another machine on the network.

The next step in my development will now be to implement the HTML interface that will read the data and pan a sound based on the given position.


Wednesday, September 25, 2013

Kinect Sample Source

I am tasked with cracking open the skeleton tracking source code and finding a way to read the data points that represent the skeleton. In particular, my end goal is to pan audio from left to right depending on the position of a given hand, say the left one.

Perhaps its my unfamiliarity with Visual Studio 2012 but upon cracking open the source code to SkeletonBasics from Microsoft Developer Tools, I do not see any C# code. The project simply contains an App.xaml and MainWindow.axml file.

Ah, apparently the .axml files that apperead in Visual Studio's Solution Explorer have expanding arrows next to them and once expanded, each axml file has an appropriate .cs source. It looks like the implementation lies here.

After figuring out where the code was, I dug in and started sifting through it to find where exactly the data for the left hand was coming from. I found the instance of the class it was coming from and modified the skeleton tracking sample to display the X Y Z of what is considered the "left wrist". The coordinates are on a range of -1 to 1 and are extremely sensitive to movement. Even after rounding to two decimal places the numbers still fluctuate even when I am sitting still.

Next, I was going to look at the sample "Webserver Basics - WPF" which according to the documentation here, sounded exactly what I was looking for in terms of pairing with the browser. Unfortunately, as soon as I tried to run the sample I got a lovely message that it requires Windows 8 because the System.Net.WebSocket class is "not implemented in your current operating system version". Further research led me to this page which apparently states that abstract classes exist for previous operating systems versions to allow a different implementation of WebSockets to be plugged in.

This might be something to consider but other than that I'll have to take a different approach to leverage the Kinect through a web browser.