Showing posts with label diss. Show all posts
Showing posts with label diss. Show all posts

Thursday, 21 May 2015

Done Done

Well 4 years later and I am finally finsihed with university. I have had a wonderful time at Aberystwtyth and I hope to not be away from it for too long. I am still waiting on results but hopefully I will get at least a 2:1(a 1st if I am lucky).

So what am I up to now, well, first priority is to get a job. I have spoken to a few people and got a few interviews lined up, so we shall see how they go. To try and keep on top of things I am working on a few projects, both hardware and software. As well as that I have a charity raft race this coming Sunday at Glengower in Aberystwyth so that should be good. Got a few friends coming up to take part as well.

Other things that are going on is my YouTube channel is back to being updated and you can have a look at what I have going on here feel free to like my videos or if you feel up to it even subscribe.

I should hopefully get back to updating my blog a bit more as well and might even start doing some Vlogs on YouTube so look out for them.


Monday, 13 April 2015

People are objects

Well I am back and I have been working hard on my dissertation. I have gotten the people detected to be stored as objects after a bit of an issue. Never having done much Python before this dissertation and no object orientated Python at that, I can up on a few issues. My first was treating it like Java and putting global variables at the top of the class file. This I later found out that these were shared by all the objects, so when modifying one object global variable I modified them all. This was resolved by initializing them in the __init__ which is the first method that is run when the object is created. This involved me putting self.var_name = value. The self lets it know that that variable belongs to only that instance of the object. This link really helped me, especially 9.3.5.

The second issue was checking all the objects to see if they were a new person or the same person object and just needed location updating. This was partially resolved by checking the location of the new detections and comparing them to the location of the people objects. This will later incorporate the Kalman filter, and will target in a set direction making it a more accurate verifier.

As can be seen here the Kalman filter is not starting at the centre of the people which needs to be worked on. This will make initial tracking more accurate rather than having to wait for it to align.

The blue Xs are the start location of the Kalman filter



My next lot of tasks to do is to add in the mean shift and cam shift to improve the tracking. This will help determine if there is still a person at a location as well as helping determine if new detections are accurate. As well as this I need to track the location of people exiting the scene. If they are disappearing along the edge of the scene then it can be assumed that they wandered off screen. If they disappear on screen or are not detected then you can increase the likely hood of something such as a door existing at that location. The same can be assumed if reappearing at the same location.

Another possibility that could be happening is that there is an obstacle obscuring the view of the person for an extended period of time.

Friday, 27 March 2015

Kalman away

Well took long enough but I have a Kalman filter working. Thanks to help and code from my dissertation supervisor Hannah Dee. I have managed to tweak her code that uses Kalman2d which can be found here to fit in with my scenario. Example of what it is currently doing can be seen below.

Here shows the Kalman filter in action(the green line) however at this point it jumps between the two people in the box currently


I have added some basic checking for the HoG(Histogram of orientated gradients) and Kalman filter so that if a HoG is not detected the bounding box is drawn in the last place it was located. If it then re appears the Kalman filter then is updated with the new location. This is done by searching with a buffer with in the range of last HoG detection, this is assuming the person has not travelled far and there is no one else nearer.

The issues I am having with two people coming near to each other that the tracker swaps over is not an issue. The reason being as if two people come close to each other and one gets the tracker from the other then the other one will get their tracker. One way I can mitigate this is to start the search in the direction the person was travelling so that way the person is more likely to get their tracker back.

I am currently only tracking one person but now I have a goal I feel I can really push this through. My goal is to build a person object that will keep track of locations to help mitigate occlusions and count them. This will be using HoG detector as the initialiser, mean shift, Kalman filter and cam shift to help determine location helping keep a more accurate track of people. Background subtraction will be used as evidence to show that something is moving here chances are someone or something is here.

I am hoping to be able to determine in scene exits by determining if a number of person detectors go dead at the same location. If they are predicted to go off scene then that is a normal exit but can monitor exits to focus detectors more.



Any way this will be my last blog post for a while as I am off to sunny Lloret de Mar for my first vacation abroad with dance sport. Should be a good time sportsvest who are organising it have made it look like it is going to be an epic time. So speak in a week and a bit.

Lloret de Mar




Thursday, 19 March 2015

Points, assumptions and more graphs!

Ok so it has been a busy week so here comes a list of what I have been up to with my dissertation. Don't worry there are pictures.

First thing I got working was MoG(Mixture of Gaussian) background subtraction using this page on OpenCV. This is to allow me to detect movement within the image that can help determine the likeliness of a person existing at that point. I have also implemented moving average background subtraction from this helpful site. I will be creating a test to try different values for moving average background subtraction to see which one is the best to use.

I have also implemented Cam shift which is starts off by doing a mean shift but then updates the size of the window and calculates the best rotation of a fitting ellipse to it. Then re applies the mean shift with the new scaled search window and previous window position. This is then repeated until the desired accuracy is met. (Paraphrased from here).


I am dealing with the problem of people being obscured by using a Kalman filter to determine the direction and velocity of people. From this I can determine if they are going to end up off screen or if they are still with in the scene but obscured. The issue may arise if there is an exit on screen that the person can disappear through. This issue can be dealt with by assuming that if the person isn't detected with in a number of frames then they are likely no longer in the scene. I won't be modelling the scene as that is a whole other project.


I have also made some more wonderful graphs to show the number of correct detections that each setting for the HoG detector provides. These can be seen below

Patterns include: Multiples of 4 having 0 detections and multiples of 6 having highest number of detections.

Steady drop of accuracy after scale setting of 1.01

This peaks at 3 but may go back up for higher numbers



My code is getting a little messier than I would like it to be. For this reason I will be going over my code and modularising it and doing some more testing for the new code I have done.


I am working on a points based system that will give points to the more accurate detections that are all consistent. This will use mean shift, cam shift and hog detectors to build up a more accurate person detector. I will create an array of ROIs(regions of interest) for each person depending on the score they get to build up what to look for.


I have encountered some problems along the way. Some of these include contours from background subtraction overlapping with other HoG boundaries meaning false positive when looking for movement in a HoG detection. Another issue is a 'non convex optimisation problem' with trying to search through the most accurate window stride for HoG detector. This is trying to avoid searching through all the possible combinations by going in the direction of the best detector settings in chunks.


An interesting point that was raised in a dissertation meeting was, should you count people that are occluded. If a person was counting number of people in a scene and a person was occluded that person would not be counted. There may be times when you want to know if a person is still in a scene even when obscured. An example would be when a place is limited to a certain number of people but they will be blocked by certain elements. In this case keeping track of number of people is important.

In the case of my dissertation I will be keeping track of all people that are judged to still be with in the scene.







Tuesday, 10 March 2015

The ups and down of graphs!!!!

After having some issues I managed to get my graphs for the different settings for the HoG detector and this is what they look like.

This is the number of correct detections for a few HoG settings

This is the number of offsets from the Ground truth for a few HoG Settings   


This is only based upon a few frame so that it did not take forever to complete. The graphs as well only contain a select few HoG settings, otherwise it would be to busy and un readable.

I am currently working on trying to try different values between bigger numbers. This will allow me to try more combinations without trying all the combinations between.

Monday, 9 March 2015

The mean shift, the background subtraction and lots of time

Over the past week and a bit I have had to re create my HoG(Histogram of orientated gradients) as I discovered I had the wrong values in it. This is currently running and will hopefully finish soon as it has been running for two days now. As well as this I have implemented mean shift to track people over a number of frames and looked at background subtraction to increase accuracy of detectors.


The various combinations of HoG detector settings come to a rough total of 20000 different combinations. These include settings such as scale between 1 and 1.1, window stride between 0 and 9 and scaling between 0 and 32. This will then be used to determine the best and most accurate settings to use for the HoG detector increasing the overall performance. An issue I had was that they were not in the correct order so I had to write a quick script to go through the lines and sort them out. Below you can see the output from the 20000 lines of combinations and results of the HoG detector on 8 test images.

As it shows best values and how many were correct and what the total offset is for that setting
This value takes a while to process an image so it may be worth searching different combinations in between. As currently, for padding which takes two numbers, I use the same one twice. So padding(6,6) rather than padding(4,8), the same can be said for window stride. This could be faster or result in a more accurate detector. This will be ran at a later date.

As well as using HoG I am also using mean shift to track the movement of people. This will be helpful in counting people when they are occluded and then picking them back up when they come into focus. An example of current progress can be seen below.

Mean shift over 10 frames trying to track top and bottom
As you can see the boxes are not that accurate or tight round the people so the tracking is not as good as I would like it to be. To try and fix this and deal with false positives like the window at the top of the image I have been looking at background subtraction. This will allow me to ignore the parts that are not moving such as the window and focus on the more likely areas to contain people. It may just still be a cat.


To also keep track of people and make sure the mean shift does not go to far away from the people I will be re checking the area to find people are still there using the HoG detector. This will be less cpu intensive as we can assume that the people, if still are in frame, will be near by so no need to check the entire image.


More graph to follow when I get them working properly :)

Wednesday, 11 February 2015

That'ssss some very nice Python code there...

Well it has been a busy week of work on my dissertation and so with out messing around lets dive straight into what I have been up to.

Firstly I have handed in my project specification in on Friday and got feedback on The following Tuesday. The feedback was very helpful and helped clear up some thing such as focusing me more on target rather than trying to solve everything. As a result of this feedback my goals are even more clear now and they are to work on counting people in a crowd and determining if the crowd is calm or not. Other areas for improvement include my bibliography which did not have all the relevant information on some of the references.

As well as this I have been carrying on my reading and spent a long time trying to get OpenCV installed for Python on Ubuntu 14.04. I got shown a nice way of installing OpenCV via pip which is a nice way of installing stuff for Python. It installed but when I tried to run any test code with windows it was throwing dependency issues. So I looked around and found this nice tutorial which helped resolve some of my issues but still had an issue with gtk. This answer managed to fix the issue and I was up and running. I have also been learning a lot about Python such as how to be modular and returning tuples, example code below.


def tuple_return_function():
        return ("10", "20")

def function_name();
        tuple1, tuple2 = tuple_return_function()

        print("Tuple1 is " + tuple1 + " Tuple2 is " + tuple2)
       # Outputs: 'Tuple1 is 10 Tuple2 is 20'


Once I was up and running I got to work trying different OpenCV implemented methods using this site. As such I got some nice SIFT and SURF and a interactive foreground extraction using GrabCut algorithm. For SIFT and SURF I also made it loop so it only found a specific number of points so wouldn't detect everything. These helped me understand the basic functionality such as loading in images, copying them and converting them to grey and other basic stuff. You can see some examples below.
SURF - Before and after


GrabCut - Before and after

SIFT - Before and after


I have also started working on some code that will be used in the final system. Currently I have a Python file that is used for reading in image files from a directory as well as a basic implementation of Viola-Jones. Once I am happy with my implementation I will then begin writing tests for both the reading in image files and the Viola-Jones. I will be starting on implementing histograms of oriented gradients for human detection next after the tests are implemented.


I am having some issues with the Viola-Jones implementation for detecting faces in a crowd and it isn't that the faces are obstructed. It seems to be that the faces are too far away to be picked up. I am having a look at this and seeing if I can tweak some settings.

Sunday, 1 February 2015

I predict a riot

So finally decided on what my dissertation is going to focus on and, drum roll please, it is crowd behaviour. Specifically trying to count out when a crowd is likely to form, how many people are in a crowd, why they are they are in a crowd and dispersal patterns of the crowd. As well as this I will have to take into account the ethics of analysing people in crowds.

The main goals of this is to be able to determine at minimum number/groups of people in crowds. From this I can try to ascertain the situation and evaluate possible outcomes. It would be nice to be able to accurately count the people in a crowd however there are issues with trying to count people close together. If they are too close then an issue arises with counting many people as one person, as well as this the faces may not all be visible or clear enough to use facial recognition or just obstructed. There has been research into counting people in crowds or mapping high density crowds but I have not come across anything that is accurate and versatile between scenarios. The research linked for 'counting people in crowds' has a high success rate of more than 96% but the data set is only three videos. With regards to the link provided in the 'mapping high density crowds' they track motion against a certain threshold and if not over the threshold then it is considered static. This raises the issue of people who are not moving fast enough which can happen in over crowded areas would be considered background. On the other hand things, such as large animals, that move at the threshold could be used in the generation of the map thus making it not accurate. Being able to determine when a crowd is about to form would help determine things such as when a riot or a big event is about to happen( discussed in more detail later).

Below is an image of where people counting could be useful to make sure there is no over crowding. As well as this it could be used to determine best flow of traffic or if something abnormal is happening. *




Within crowds there are specific behaviours that can looked out for to help identify individual people or to try and work out what is going on in the scene. When walking down the street and another person is heading towards you the closer they get the more you move to the side to pass them as explained here. When people are coming from different angles but heading in the same direction they tend to merge in to a single flow of traffic. There is however, no specific detailed definition of a crowd but is defined as 'a large number of people gathered together in a disorganized or unruly way'. A few definitions do exist and share common specifics such as 'conceptualising a crowd as a sizeable number of people gathered at a specific location for a measurable time period, with common goals and displaying common behaviours' which is a extract from this online pdf. This document also goes into more detail about what is expected from a crowd. There is however issues with things such as determining when a group of people are considered a crowd, such as number of people and time spent together/at a location. For these reasons a set definition of a crowd must be made to allow a system to appropriately determine if there is a crowd or if a crowd is likely to form. As well as defining a crowd, definitions of crowds in different situations from data sets will allow the system to more clearly work out what is going on.

The accuracy of this is likely to be incrementally lower as the crowds get bigger as it will be harder to count people and monitoring the flow will become process heavy. As such focusing on each part such as the counting will allow more targeted results with hopefully higher accuracy. This could be useful though for predicting violence in crowds in which some work has been done here or even counting people in areas to avoid overcrowding and injuries.

Below image shows how the faces are not always showing on cameras, this makes it more difficult to count people using facial recognition. *




Now comes the fun part, ethics and what is ok to use, after all we will be looking at humans and their behaviour.The first thing we need to make sure is that the people who are on the video are ok to have themselves used for research purposes and their privacy is protected. Some questions need to be asked such as ones raised in this papers abstract. Questions such as 'Under what conditions should video be presented and to which audiences'. Videos that are used should have the consent of the people in it and should only be used for the purposes they are made for. As well as this there should be no attempt to try and identify persons within the videos unless that is the reason for the videos and you have express permission from all persons involved.


This is just a little overview of the parts of what needs to be discussed and will be discussed in more detail over the coming weeks.





Tuesday, 27 January 2015

The begginning of the end!

Well I am coming closer and closer to the end of university and as such that means it is dissertation time. I am going to keep a diary of what I have been up to with regards to my dissertation and keep up to date weekly. You can follow along and see what I go through and offer help or even learn some stuff.

First things first what am I doing?

Well it is called scene interpretation, which basically means I will take a video of something such as a street view that is publicly accessible and analyse the environment or people. With the people aspect there is some ethical concerns but thankfully there is already large data set that are available to use here. My first step though with regards to my dissertation project is to decide what aspect I am to focus on, whether people/crowd analysis or environment analysis. This will then be the core focus of my work and will stop me branching out to far and creating to much work for myself.


As well as deciding what to focus on I also am having a look at openCV in Python. I have found this link to a good tutorial, but I dare say I will be using multiple resources for learning as I go. I have used Python before but never for a project so this could be quite interesting. There will also be some background reading done on information about scene analysis to be able to use some examples or learn from.


My first goal is to get a system up and running that can do basic processing of videos and get some kind of recognition going on it. I have a meeting later this week so I can discuss in more detail where to go from there.