 |
 |
|
 |
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
http://www.eecs.harvard.edu/~cduan/technical/git/
Having spoken with half a dozen people who said "I hate git, it's so
confusing" and then after showing them this they go "wow, that's really
easy", I figured it might be worthwhile to show people this. :-)
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Le 19/04/2011 20:25, Darren New nous fit lire :
> http://www.eecs.harvard.edu/~cduan/technical/git/
>
> Having spoken with half a dozen people who said "I hate git, it's so
> confusing" and then after showing them this they go "wow, that's really
> easy", I figured it might be worthwhile to show people this. :-)
>
Let's start a holy troll war. best OS or browser is so 20th century.
I prefer mercurial. ;-)
Oh, there is a false assertion about rebasing being unique to git on the
last page.
And doing graph in ascii art... ouch on the web!
(really, graphviz is not that hard to learn)
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 19/04/2011 08:45 PM, Le_Forgeron wrote:
> I prefer mercurial. ;-)
I prefer Darcs. Then again, I haven't tried any other [distributed] VC
system, so...
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/19/2011 12:45, Le_Forgeron wrote:
> Let's start a holy troll war. best OS or browser is so 20th century.
I honestly wasn't recommending git.
However, that said, and with the understanding that I haven't looked at the
other DVCSs, I think git has a very interesting layered approach to the problem.
First you have the virtual file system that is the repository.
Then on top you have actual commits and such, almost completely independent
of the file system model.
Then on top of that you have stuff like packs and transfer protocols and
such, again almost completely independent of the file system and commit model.
Kind of cool, in much the same way that Second Life has a kind of cool
business model regardless of how well they actually implemented the code.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 19/04/2011 22:05, Darren New wrote:
> I honestly wasn't recommending git.
>
> However, that said, and with the understanding that I haven't looked at
> the other DVCSs, I think git has a very interesting layered approach to
> the problem.
OK, I just sat down and read it.
Yeah, it's fairly straight-forward. But I can see how if you don't
understand the way it works, there's huge potential to *seriously*
confuse yourself!
I had assumed that all DVCSs were the same, but I now see that at least
Git and Darcs use fundamentally different models.
Fundamentally, any version control system tracks changes to files. What
Git appears to be doing is something similar to RCS or CVS, where each
file goes through a strictly sequential series of "versions", and one
version comes "before" or "after" another. It seems that each commit
stores the complete state of the entire repository. (Presumably as a
diff relative to the previous commit, but still logically it's a
snapshot of everything.) Git then uses "heads" to point to the most
recent commit in each branch, or to other interesting points in the history.
Darcs works completely differently. Darcs doesn't track sequential file
versions, it tracks change-sets. It defines a "change-set algebra" where
unrelated changes to the same file are independent, and can be applied
or reverted independently. Changes to the same part of a file are not
independent, and can only be applied in sequence. But unrelated changes
are... unrelated.
As an example, if I add some comments to file X and commit that, and
then I add a new function to file X and commit that too, I get two
change-sets. I can revert adding the comments but still keep the new
function, even though the latter change happened *before* the former.
As far as I can tell, Git would require me to create a branch where I
add the comments, and another branch where I add the new code, and then
merge them back into the main branch, hoping that I don't get any
conflicts. To me, this seems like a lot more work and a lot more
conceptual overhead.
It also seems that when you ask Git to perform a commit, you have to
tell it which files to record changes for. With Darcs, you tell it which
files to monitor, and when you ask to commit it detects what's changed
and interactively asks you which differences to include and which ones
not to include. E.g., I could add comments to file X, add a new
function, do a commit and interactively split the modifications into two
separate commits. (It doesn't /always/ work, of course, but mostly it does.)
I wonder how well the illusion of one single sequence of file versions
works when you have multiple people editing the file in parallel. I
would imagine the Darcs model works better, because it doesn't try to
pretend that file X looked like this, and then this, and then this. It
just records what edits happened, without recording their relative
ordering [except where they affect the same lines of code]. Git, on the
other hand, appears to be trying to track what every file in the entire
repository looked like in every individual commit object.
With Darcs, you make a branch by copying the repository. That's it.
(Although there is an option to build a bunch of symlinks for you
instead of just copying, to save a bit of disk space. Presumably only on
POSIX platforms...) You can email individual change-sets around, and
this works. Getting somebody else's changes just copies all change-sets
from their repository into yours. You can then resolve any conflicts.
Still, I've used Darcs quite a bit, and I've never once used Git, so I'm
not really qualified to say how well it works in practise.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 20/04/2011 10:44, Invisible wrote:
> I had assumed that all DVCSs were the same, but I now see that at least
> Git and Darcs use fundamentally different models.
>
> Fundamentally, any version control system tracks changes to files. What
> Git appears to be doing is something similar to RCS or CVS, where each
> file goes through a strictly sequential series of "versions", and one
> version comes "before" or "after" another. It seems that each commit
> stores the complete state of the entire repository. (Presumably as a
> diff relative to the previous commit, but still logically it's a
> snapshot of everything.) Git then uses "heads" to point to the most
> recent commit in each branch, or to other interesting points in the
> history.
>
> Darcs works completely differently. Darcs doesn't track sequential file
> versions, it tracks change-sets. It defines a "change-set algebra" where
> unrelated changes to the same file are independent, and can be applied
> or reverted independently. Changes to the same part of a file are not
> independent, and can only be applied in sequence. But unrelated changes
> are... unrelated.
It appears Mercurial works the same was as Git:
http://mercurial.selenic.com/wiki/UnderstandingMercurial
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
One thing which apparently can't be done with Darcs, Git or Mercurial is
managing multiple repositories at once.
For example, consider Glasgow Haskell Compiler, which is written in
Haskell itself. The compiler, interpreter and run-time system are one
repo, and the various standard libraries it requires are independent
projects with their own repos, but GHC mirrors a copy of each, lagging
behind the upstream slightly. The main GHC repo contains a special Bash
script to do things like update all the sub-repos automatically. The
existence of this script tells you that there's functionality that Darcs
itself is failing to provide. (And Git and Mercurial, apparently.)
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 19/04/2011 19:25, Darren New wrote:
> Having spoken with half a dozen people who said "I hate git, it's so
> confusing" and then after showing them this they go "wow, that's really
> easy", I figured it might be worthwhile to show people this. :-)
As an aside, I do envy you to some extent for actually knowing people
IRL who know how a computer works...
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/20/2011 2:44, Invisible wrote:
> I had assumed that all DVCSs were the same, but I now see that at least Git
> and Darcs use fundamentally different models.
GIT stores files, and deduces change sets from those files.
Mercurial stores change sets, and deduces files from those changesets.
> Fundamentally, any version control system tracks changes to files.
No, git actually tracks entire files. Every time you check something in, any
file you add to staging (i.e., any file with any change) is stored in its
entirety in the repository. And, technically, *every* file in the entire
repository is stored, but if you haven't changed it, it has the same sha-1
name, so it doesn't physically get copied.
Only later do you compress the repository into diffs. But a file can be
stored as a diff from a different file it was never related to, that was
created whole-hat *after* you checked in the file it's a diff from.
> What Git
> appears to be doing is something similar to RCS or CVS, where each file goes
> through a strictly sequential series of "versions", and one version comes
> "before" or "after" another.
Nope. Not even close. git is storing the entire repository on every commit.
That's the completely boggling idea there.
> It seems that each commit stores the complete
> state of the entire repository.
Yes. So there's no "before" or "after" for files, technically. There's
before or after for repositories.
> (Presumably as a diff relative to the previous commit,
Nope. It stores the entire repository. Now, if you don't change a file, it
hashes to the same value, and hence doesn't need to get stored again. But
the entire file is put into the repository.
That's why git doesn't have a "rename" command. That would imply you're
storing something other than files in the repository. git looks at the same
contents disappearing from one part of the directory structure and showing
up in another and says "Gee, that must have been a rename." If there are
minor changes between what disappeared on this commit and what showed up
somewhere else on that commit, git says "there's a 97% probability this was
a renamed file."
When you do a "git gc" to collect garbage, it then looks for a good way to
generate diffs between files, and it looks through all the files (not just
ancestors) to find good versions to diff from, and it stores the diffs. But
that's just an storage optmization, just like the fact that it gzips the
files in the repository is a storage optimization.
If I checked in a configuration file with my name in it, and 2 days later
deleted it, and then you checked in a similar configuration file with your
name in it, and then we did a "git gc", it's entirely possible (likely,
even) that my file would be stored as a diff from your file.
In contrast, it seem Mercurial actually stores a tree of changes.
The only reason git uses a pointer to earlier commits is when you merge
things, you don't want to apply changes you already applied in an earlier
merge.
For example, when you merge branch A into branch B, git finds the common
ancestor (i.e., the point at which you separated B from A in the first
place), then *generates* the diff between what's now A and where they
branched, then *applies* that diff to B, and leaves it ready for a commit.
There's no diff stored before you type "git merge" or after it returns from
the command line.
Reverting a commit involves either just throwing away the new version, or
generating a reverse diff and applying it.
That's the thing that throws people about git, apparently. They're thinking
in terms of diffs, when it's easiest to understand if you just think about
it in terms of file contents.
> Git then uses "heads" to point to the most recent commit in each branch, or to
> other interesting points in the history.
Right.
> Darcs works completely differently. Darcs doesn't track sequential file
> versions, it tracks change-sets. It defines a "change-set algebra" where
> unrelated changes to the same file are independent, and can be applied or
> reverted independently. Changes to the same part of a file are not
> independent, and can only be applied in sequence. But unrelated changes
> are... unrelated.
Yeah, from the little I read about it, Darcs is another one of those
"interesting" ideas. An actual mathematical system for defining a
repository, like relational algebra did for databases.
> As an example, if I add some comments to file X and commit that, and then I
> add a new function to file X and commit that too, I get two change-sets. I
> can revert adding the comments but still keep the new function, even though
> the latter change happened *before* the former.
>
> As far as I can tell, Git would require me to create a branch where I add
> the comments, and another branch where I add the new code, and then merge
> them back into the main branch, hoping that I don't get any conflicts. To
> me, this seems like a lot more work and a lot more conceptual overhead.
Nah. That's only if you want to have both at once working in parallel. That
is, if you want one version with the comments but no function, and another
version with the function but no comments, that's trivial in git. If you
then want to combine them into a third version that has both comments and
function, then you merge, which is also trivial unless you changed the same
lines in both places. (I.e., it's as trivial as any other diff-patch based
merge.)
> It also seems that when you ask Git to perform a commit, you have to tell it
> which files to record changes for.
Sure. But you can say "add all changes" trivially. Or you can use
interactive tools to commit just bits and pieces of this and that.
Basically, there's a "staging" area where you build a new copy of the
repository by including files and directories that are different from what's
out there already. Then you add those new files and directories to the
repository and point a commit to the new top-level directory.
> With Darcs, you tell it which files to
> monitor, and when you ask to commit it detects what's changed and
> interactively asks you which differences to include and which ones not to
> include. E.g., I could add comments to file X, add a new function, do a
> commit and interactively split the modifications into two separate commits.
> (It doesn't /always/ work, of course, but mostly it does.)
This is trivial with GIT. I do it all the time. I'll be adding a new
function, and while testing, realize there's a bug in some other function.
So when everything works again, I'll do two commits, staging just particular
hunks (in the diff sense of the word) and do two commits, one for the bugfix
and one for the new change.
> I wonder how well the illusion of one single sequence of file versions works
> when you have multiple people editing the file in parallel.
There's no single sequence of file versions. Every file is a new version.
Given that it's the repository format used by Linux developers, I think it's
safe to say it works adequately for multiple people editing the file in
parallel.
> I would imagine
> the Darcs model works better, because it doesn't try to pretend that file X
> looked like this, and then this, and then this.
Neither does git.
> It just records what edits
> happened, without recording their relative ordering [except where they
> affect the same lines of code]. Git, on the other hand, appears to be trying
> to track what every file in the entire repository looked like in every
> individual commit object.
Yes, but since you have them all, you can recreate the diffs between any two
versions whenever you want.
> With Darcs, you make a branch by copying the repository. That's it.
In git you make a branch by saying "make a new branch." That's what boggled
me about mercurial. Really, I need multiple repositories to let me have a
stable version and a development version?
> (Although there is an option to build a bunch of symlinks for you instead of
> just copying, to save a bit of disk space. Presumably only on POSIX
> platforms...) You can email individual change-sets around, and this works.
> Getting somebody else's changes just copies all change-sets from their
> repository into yours. You can then resolve any conflicts.
git is exactly the same, except it copies files instead of changes. When you
say "pull from that repository", it finds each "head" (i.e., branch tip) and
then recurses through the data structures pulling anything reachable from
them. Since they're all named after the hashes, if you already have a file
of the same name, you don't need to copy it. Then it stores the heads for
the remote repository in a different place of the namespace tree than your
own heads.
But, really, the only things in the git repository are files (blobs),
directories (trees), commits (a log message pointing to a tree and maybe
other commits), and tags (a log message pointing to a commit, possibly
pgp-signed), and then a bunch of names for specific commits or tags or trees
(i.e., branch names). There are no diffs. There is no history. There are no
users.
If you want the log history, you follow pointers from one of the heads thru
the different commit objects. If you want to see what changed, you run diff
on the two files you want to know what changed between. If you want a new
branch, you just modify some files, store them, and point a different name
at the new commit.
If you want to merge someone's repository into yours, you simply copy from
them any files or names that they have that you don't, and you're done.
You're merged. Now if you want to incorporate their changes into your work,
you generate a diff between their latest version and some earlier version,
and apply that diff to your latest version, and you're merged.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/20/2011 3:08, Invisible wrote:
> One thing which apparently can't be done with Darcs, Git or Mercurial is
> managing multiple repositories at once.
git can do this in a sorta half-assed way. Mercurial apparently can too.
> you that there's functionality that Darcs itself is failing to provide. (And
> Git and Mercurial, apparently.)
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/20/2011 8:38, Darren New wrote:
> On 4/20/2011 3:08, Invisible wrote:
>> One thing which apparently can't be done with Darcs, Git or Mercurial is
>> managing multiple repositories at once.
>
> git can do this in a sorta half-assed way. Mercurial apparently can too.
Wrong button. :-)
http://ssteiner.wordpress.com/2008/12/30/git-subprojects/
The fundamental problem is that all the DVCS systems tend to use crypto
hashes to identify things. So in git, for example, when you create a
sub-project, the parent project records where the repository is and what
commit to check out. If you change the sub-project, the old commit is still
there; you have just added to it. So the parent project is still going to
get the old version of the subproject until you tell the parent project
"hey, go update your pointer to the sub project to be commit 01A73F9E."
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 20/04/2011 16:40, Darren New wrote:
> The fundamental problem is that all the DVCS systems tend to use crypto
> hashes to identify things. So in git, for example, when you create a
> sub-project, the parent project records where the repository is and what
> commit to check out. If you change the sub-project, the old commit is
> still there; you have just added to it. So the parent project is still
> going to get the old version of the subproject until you tell the parent
> project "hey, go update your pointer to the sub project to be commit
> 01A73F9E."
The problem Git seems to have is that it uses heads to keep track of
things. Delete the head and the corresponding commit drops off the face
of the Earth.
Darcs manages a set [as in set theory] of changes. You don't need to
keep updating a "pointer" to point to the latest one or anything. I'd be
surprised if no over VCS has thought of this.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
>> I had assumed that all DVCSs were the same, but I now see that at
>> least Git
>> and Darcs use fundamentally different models.
>
> GIT stores files, and deduces change sets from those files.
>
> Mercurial stores change sets, and deduces files from those changesets.
Doesn't appear to me that that's what happens, from what little
Mercurial documentation I've read.
>> Fundamentally, any version control system tracks changes to files.
>
> No, git actually tracks entire files.
Fundamentally, VCS are about tracking changes. Git might *implement*
that by storing the entire file, but *logically* what you're trying to
do is keep track of what you changed.
> And, technically, *every* file in the entire repository is stored.
Yes, I gradually game to that realisation. Git is managing the entire
repo as a strictly linear sequence of unumbered versions. (Until you
explicitly create branches, anyway.)
>> (Presumably as a diff relative to the previous commit,
>
> Nope. It stores the entire repository. Now, if you don't change a file,
> it hashes to the same value, and hence doesn't need to get stored again.
> But the entire file is put into the repository.
How odd... Still, if you're not worried about the internal
implementation, logically Git is versioning the whole repo as one unit,
and that's all you need to know.
> That's why git doesn't have a "rename" command.git looks at the
> same contents disappearing from one part of the directory structure and
> showing up in another and says "Gee, that must have been a rename." If
> there are minor changes between what disappeared on this commit and what
> showed up somewhere else on that commit, git says "there's a 97%
> probability this was a renamed file."
o_O
OK, wow. I thought having to tell Darcs when I rename stuff was
inconvenient, but this just sounds insane...
> The only reason git uses a pointer to earlier commits is when you merge
> things, you don't want to apply changes you already applied in an
> earlier merge.
And here I was thinking it was so you can revert to earlier versions if
you want. You know - the entire purpose for a VCS to exist in the first
place? ;-)
> Yeah, from the little I read about it, Darcs is another one of those
> "interesting" ideas. An actual mathematical system for defining a
> repository, like relational algebra did for databases.
It sounds simple enough. If this change affects line X and that change
affects line Y, they are independent.
Ah, but wait. What if some change adds or removes lines? If change X
adds a new line between lines 50 and 51 then change Y no longer affects
line 150, it now affects line 151. But X and Y are still independent.
There's more to it than meets the eye. Of course, if you just want to
*use* Darcs, you just edit stuff and it "just works".
>> As far as I can tell, Git would require me to create a branch where I add
>> the comments, and another branch where I add the new code, and then merge
>> them back into the main branch, hoping that I don't get any conflicts. To
>> me, this seems like a lot more work and a lot more conceptual overhead.
>
> Nah. That's only if you want to have both at once working in parallel.
Isn't "working on both at once" kind of the entire point of distributed
version control?
> That is, if you want one version with the comments but no function, and
> another version with the function but no comments, that's trivial in
> git. If you then want to combine them into a third version that has both
> comments and function, then you merge, which is also trivial unless you
> changed the same lines in both places. (I.e., it's as trivial as any
> other diff-patch based merge.)
And if you merge the comments branch into the main branch, and then
somebody adds more stuff to the comments branch, then what?
>> It also seems that when you ask Git to perform a commit, you have to
>> tell it which files to record changes for.
>
> Sure. But you can say "add all changes" trivially. Or you can use
> interactive tools to commit just bits and pieces of this and that.
With Darcs, I tell it what files to watch, and then when I've finished
editing stuff, I say "record this" and it shows me every modified line
of every file and asks which modifications to keep. Git doesn't support
recording half a file modification, and doesn't even figure out which
files changed.
> This is trivial with GIT. I do it all the time. I'll be adding a new
> function, and while testing, realize there's a bug in some other
> function. So when everything works again, I'll do two commits, staging
> just particular hunks (in the diff sense of the word) and do two
> commits, one for the bugfix and one for the new change.
Given that Git can only record the new file or the old one, how is that
possible?
>> I wonder how well the illusion of one single sequence of file versions
>> works when you have multiple people editing the file in parallel.
>
> There's no single sequence of file versions. Every file is a new version.
>
> Given that it's the repository format used by Linux developers, I think
> it's safe to say it works adequately for multiple people editing the
> file in parallel.
This boggles my mind. Apparently I /don't/ understand how Git works at
all, because the way it seems to work precludes two people touching the
same file at the same time...
>> It just records what edits
>> happened, without recording their relative ordering [except where they
>> affect the same lines of code]. Git, on the other hand, appears to be
>> trying
>> to track what every file in the entire repository looked like in every
>> individual commit object.
>
> Yes, but since you have them all, you can recreate the diffs between any
> two versions whenever you want.
That's my point. If multiple people are editing the same files, you do
*not* have all the changes.
>> You can email individual change-sets around, and this works.
>> Getting somebody else's changes just copies all change-sets from their
>> repository into yours. You can then resolve any conflicts.
>
> git is exactly the same, except it copies files instead of changes.
And the "minor detail" that if 200 people edit the same file, that's 200
separate branches which have to be manually merged back together again.
> If you want to merge someone's repository into yours, you simply copy
> from them any files or names that they have that you don't, and you're
> done. You're merged.
It would be nice if Darcs worked that way.
> Now if you want to incorporate their changes into
> your work, you generate a diff between their latest version and some
> earlier version, and apply that diff to your latest version, and you're
> merged.
What a backwards way to look at it.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/20/2011 8:49, Invisible wrote:
> The problem Git seems to have is that it uses heads to keep track of things.
> Delete the head and the corresponding commit drops off the face of the Earth.
Yes. That's why you shouldn't do that.
First, deleting a head that you can't reach from anywhere else requires you
to answer a confirmation, just like anything else. Second, the *files* are
still there. You just might not know what they're called. They're probably
still around at least a couple of weeks before git cleans them up. I.e.,
there are well-documented ways to recover from this if you do it accidentally.
On the other hand, if you work on something and decide it wasn't a good
idea, you can delete the branch and no harm no done. Darcs apparently
requires you to copy the entire repository before you even *start* making
changes if you want to recover.
> Darcs manages a set [as in set theory] of changes. You don't need to keep
> updating a "pointer" to point to the latest one or anything. I'd be
> surprised if no over VCS has thought of this.
But that's exactly why you need to start a new repository if you want a new
branch. If you clone a repository in Darcs, make a bunch of changes, then
accidentally delete the repository, you're in even worse shape than if you
delete a branch in git.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/20/2011 9:03, Invisible wrote:
> Doesn't appear to me that that's what happens, from what little Mercurial
> documentation I've read.
I don't know. All the mercurial documentation I've read talks about change sets.
> Fundamentally, VCS are about tracking changes.
Fundamentally, they're about controlling versions. :-)
> Git might *implement* that by
> storing the entire file, but *logically* what you're trying to do is keep
> track of what you changed.
I think it depends. If I want version 1.0 that was released, I don't really
care what changed to get there. I want that version. There's uses for
changes, and uses for storing what you stored.
The advantage of storing what's actually there is you can write all kinds of
better tools to tell you the differences. For example, you can diff any two
files to get a compressed copy of the file. If I change a file to include an
additional 500 lines, then change it to delete 498 of them, the third
version is going to get stored as a 2-line diff from the first version, not
a 498-line diff from the second version.
Basically, you're storing absolutes and deducing differences, rather than
storing differences and deducing absolutes. That means when you want to know
what changed between release candidate 3.5RC2 and the version 4.2 that Fred
compiled over on *his* machine, you can just compare the two. You don't have
to reconstruct anything first.
> Yes, I gradually game to that realisation. Git is managing the entire repo
> as a strictly linear sequence of unumbered versions. (Until you explicitly
> create branches, anyway.)
Or until you clone it, yes.
>>> (Presumably as a diff relative to the previous commit,
>>
>> Nope. It stores the entire repository. Now, if you don't change a file,
>> it hashes to the same value, and hence doesn't need to get stored again.
>> But the entire file is put into the repository.
>
> How odd... Still, if you're not worried about the internal implementation,
> logically Git is versioning the whole repo as one unit, and that's all you
> need to know.
Yes, basically. That's why you can sign just the tag blob and be sure you've
signed every file that that tag refers to.
> OK, wow. I thought having to tell Darcs when I rename stuff was
> inconvenient, but this just sounds insane...
Why? You don't have to tell git you renamed something.
>> The only reason git uses a pointer to earlier commits is when you merge
>> things, you don't want to apply changes you already applied in an
>> earlier merge.
>
> And here I was thinking it was so you can revert to earlier versions if you
> want. You know - the entire purpose for a VCS to exist in the first place? ;-)
Well, yes, it gives you a way to find those commits. But in theory you could
look up any commit by hash code and say "give me that version" whether it's
earlier or later or completely unrelated to what you have now. Indeed,
that's exactly how branches work. When you start a new branch, all you're
doing is storing the commit sha-1 into a new file named after the branch.
> It sounds simple enough. If this change affects line X and that change
> affects line Y, they are independent.
Yeah, until you get binary objects in there. :-)
>>> As far as I can tell, Git would require me to create a branch where I add
>>> the comments, and another branch where I add the new code, and then merge
>>> them back into the main branch, hoping that I don't get any conflicts. To
>>> me, this seems like a lot more work and a lot more conceptual overhead.
>>
>> Nah. That's only if you want to have both at once working in parallel.
>
> Isn't "working on both at once" kind of the entire point of distributed
> version control?
Only if you want to work on both at once in the same repository.
If you're working in a different repository, you don't need to start a new
branch. Branches are nothing but names for commits. You technically never
need to use any branch at all if you want to type in a sha-1 every time.
>> That is, if you want one version with the comments but no function, and
>> another version with the function but no comments, that's trivial in
>> git. If you then want to combine them into a third version that has both
>> comments and function, then you merge, which is also trivial unless you
>> changed the same lines in both places. (I.e., it's as trivial as any
>> other diff-patch based merge.)
>
> And if you merge the comments branch into the main branch, and then somebody
> adds more stuff to the comments branch, then what?
Then you get a merge and then more changes on the comments branch. And if
you merge the comments branch *again*, *that* is when git uses the parent
pointers in the commit objects to figure out which files to diff in order to
get the patches to the parent.
A--B--C--D--E--F--G--H--I
\ | /
Q--R/--S--T/
So you started with A, changed to B, branched B and made a change to create
Q, then R. In the mean time, I changed B to be C. Now I merge your R back
to my C. This looks back, sees B is the common ancestor, so diffs R against
B and applies it to C, then creates D with C and R as parent commits. (Each
letter is a commit, which includes the state of the entire repository.)
Now you keep working on R without incorporating my B->C change, creating S
and T. I change D to include E and F. Now I merge your work again.
Git looks at F, follows it back to D, to C and R, and sees that R is a
common ancestor of both F and T. So it diffs T against R, applies those
diffs to F, and creates G. You can then delete the branch that points to T
safely without losing anything.
It's *super* straightforward to understand what merges do in git.
And if someone comes up with a better diff algorithm, no problem. The
algorithm to do the diff during a merge isn't built into the repository.
> With Darcs, I tell it what files to watch, and then when I've finished
> editing stuff, I say "record this" and it shows me every modified line of
> every file and asks which modifications to keep. Git doesn't support
> recording half a file modification,
Yes it does. Indeed, you can even go back and retroactively say "oh, those
two commits? The second one should have come first, and the first one should
be broken up into these three commits."
As I said, I do this all the time.
> and doesn't even figure out which files changed.
Yes it does. It compares the working directory against the staging directory
and the head to say "these files are changed and unstaged, those are changed
but already staged, and those are unchanged."
It's just a two-step process. You can build up the thing you want to commit,
and then finally commit it. It sounds like Darcs needs you to do that all
in one step.
Git does it the other way around. First it asks you what modified lines you
want to put in the commit (and puts them in the staging area), then it
creates the commit (based on the staging area).
>> This is trivial with GIT. I do it all the time. I'll be adding a new
>> function, and while testing, realize there's a bug in some other
>> function. So when everything works again, I'll do two commits, staging
>> just particular hunks (in the diff sense of the word) and do two
>> commits, one for the bugfix and one for the new change.
>
> Given that Git can only record the new file or the old one, how is that
> possible?
The staging area lies between the repository and the working directory. So I
check out some branch, and that copies it to the WD and maybe clears the
staging area. The staging area is basically a commit that's not yet in the
repository.
Now I make changes to the WD.
Then I use something like "git add" to add all the changes from the WD to
the staging directory. Or I use "git add -i" (or, more likely, the GUI) to
diff the WD against the staging area (or the repository), pick (say) three
of the five diff hunks, and then create a new temp file that holds the
repository with those three diff hunks applied, which I then put in the
staging area. When I have everything the way I like, I commit the change,
which copies the staging area into the repository and then adds a commit
object pointing to it.
>>> I wonder how well the illusion of one single sequence of file versions
>>> works when you have multiple people editing the file in parallel.
>>
>> There's no single sequence of file versions. Every file is a new version.
>>
>> Given that it's the repository format used by Linux developers, I think
>> it's safe to say it works adequately for multiple people editing the
>> file in parallel.
>
> This boggles my mind. Apparently I /don't/ understand how Git works at all,
> because the way it seems to work precludes two people touching the same file
> at the same time...
Sure. But you're thinking git tracks diffs. That's exactly the point. If I
change the file, and you change the file, then now there's three files. The
original, the new one I have, and the new one you have. When we go to merge
it, we create number four, which is your new one with the differences
between my version and the original applied.
It works because if there's no merge conflicts, then my diff applied to your
file and your diff applied to my file creates the same file.
>> Yes, but since you have them all, you can recreate the diffs between any
>> two versions whenever you want.
>
> That's my point. If multiple people are editing the same files, you do *not*
> have all the changes.
Well, no, obviously. Welcome to DVCS. If you don't give me your files, I
can't see them. This is true of changes you don't push in Darcs and
mercurial too. Maybe I'm misunderstanding what you're trying to say.
Darcs *is* distributed, right? If you change a file and check it into your
local repository, and I change it and check it into my local repository, I
can't see your changes and you can't see mine until we connect the
repositories again, right?
> And the "minor detail" that if 200 people edit the same file, that's 200
> separate branches which have to be manually merged back together again.
And this differs from any other VCS how?
Note that if you're trying to *push* changes to a remote repository, you
have to do it to a branch where nobody else has branched off since you did.
In other words, if I say "update my repository to the DEV branch on the
company's central reposityro", and I make changes, and someone else changes
the DEV branch to point to a later version, I can no longer push my changes
into the DEV branch. Instead, I have to fetch down the new DEV branch, merge
my changes, then push the newly merged commit back up. Look up "fast-forward
merge" in the git docs if you care.
But basically what it's saying is if you're *pushing* changes to a
repository (i.e., there's no human there checking the merges) then you can't
do a two-parent merge commit. You have to create the two-parent merge commit
on your own machine, *then* push it up to the server. Sorta.
>> If you want to merge someone's repository into yours, you simply copy
>> from them any files or names that they have that you don't, and you're
>> done. You're merged.
>
> It would be nice if Darcs worked that way.
Right. In Darcs, you have to merge all the changes. In git, you have to
merge all the changes.
>> Now if you want to incorporate their changes into
>> your work, you generate a diff between their latest version and some
>> earlier version, and apply that diff to your latest version, and you're
>> merged.
>
> What a backwards way to look at it.
Only if you're used to looking at source control as a series of diffs to
start with. But that's (A) exactly what makes git hard to understand and (B)
exactly what makes git brilliant. :-)
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/20/2011 9:03, Invisible wrote:
> Fundamentally, VCS are about tracking changes. Git might *implement* that by
> storing the entire file, but *logically* what you're trying to do is keep
> track of what you changed.
Or, as an alternate example, say you've been working and every day you
commit before lunch and you commit before you go home, even if it's not
working, just so it gets backed up. And you implement two functions, and you
write code on that, and then realize you should have put that first function
elsewhere, and you don't need the second function at all, and the other code
should be in separate objects, and etc etc etc.
And at the end of the week, you have 50 messy changes committed.
With git, you can say "OK, go diff the current version against where I
branched, and give me exactly one commit with all the changes I need." It's
trivial to do that in git and then say "now commit *that* change for
everyone else to see, and abandon all the intermediate changes."
I don't know how you'd do something like that in mercurial or darcs that
store *changes* in the repository.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 20/04/2011 18:11, Darren New wrote:
> Or, as an alternate example, say you've been working and every day you
> commit before lunch and you commit before you go home, even if it's not
> working, just so it gets backed up. And you implement two functions, and
> you write code on that, and then realize you should have put that first
> function elsewhere, and you don't need the second function at all, and
> the other code should be in separate objects, and etc etc etc.
>
> And at the end of the week, you have 50 messy changes committed.
>
> With git, you can say "OK, go diff the current version against where I
> branched, and give me exactly one commit with all the changes I need."
> It's trivial to do that in git and then say "now commit *that* change
> for everyone else to see, and abandon all the intermediate changes."
>
> I don't know how you'd do something like that in mercurial or darcs that
> store *changes* in the repository.
Assuming that your working copy matches everything Darcs has in its
history, you'd do this:
1. You "unrecord" the 50 messy commits. That doesn't do anything to your
working copy, just the history Darcs keeps.
2. You "record" a single commit. When you do this, Darcs diffs the whole
working copy against what it has in its history, and records that.
(In other words, if you add 500 lines, commit, delete 450 of those lines
commit, and then you unrecord the two commits and record a new commit,
the 450 lines that you added then deleted don't show up any more.)
Needless to say, you do *not* want to be unrecording any history which
other people have copies of. But if it's only your local repo, it's fine.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
>> The problem Git seems to have is that it uses heads to keep track of
>> things.
>> Delete the head and the corresponding commit drops off the face of the
>> Earth.
>
> Yes. That's why you shouldn't do that.
It's also why having sub-repos might be tricky. Much simpler if you
don't need to keep updating pointers.
> On the other hand, if you work on something and decide it wasn't a good
> idea, you can delete the branch and no harm no done. Darcs apparently
> requires you to copy the entire repository before you even *start*
> making changes if you want to recover.
What craziness are you speaking? If you want to go back to an older
version, you just say "take me back to an older version please". If you
don't want changes you've made, you either record commits reverting
them, or you just delete them from the history outright. That's kind of
the whole point of version control, distributed or not.
>> Darcs manages a set [as in set theory] of changes. You don't need to keep
>> updating a "pointer" to point to the latest one or anything. I'd be
>> surprised if no over VCS has thought of this.
>
> But that's exactly why you need to start a new repository if you want a
> new branch. If you clone a repository in Darcs, make a bunch of changes,
> then accidentally delete the repository, you're in even worse shape than
> if you delete a branch in git.
Well, yes, if you delete all your work, you have a problem. This isn't
unique to Darcs. I'm not seeing what your point is...
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Le 20/04/2011 19:11, Darren New a écrit :
> On 4/20/2011 9:03, Invisible wrote:
>> Fundamentally, VCS are about tracking changes. Git might *implement*
>> that by
>> storing the entire file, but *logically* what you're trying to do is keep
>> track of what you changed.
>
> Or, as an alternate example, say you've been working and every day you
> commit before lunch and you commit before you go home, even if it's not
> working, just so it gets backed up. And you implement two functions, and
> you write code on that, and then realize you should have put that first
> function elsewhere, and you don't need the second function at all, and
> the other code should be in separate objects, and etc etc etc.
>
> And at the end of the week, you have 50 messy changes committed.
Yes, but that is in your messy repository only.
"Commit often, Push when working" is a good approach with DVCS.
>
> With git, you can say "OK, go diff the current version against where I
> branched, and give me exactly one commit with all the changes I need."
> It's trivial to do that in git and then say "now commit *that* change
> for everyone else to see, and abandon all the intermediate changes."
>
> I don't know how you'd do something like that in mercurial or darcs that
> store *changes* in the repository.
>
For mercurial, there is an extension which aggregate the change-line or
even a cloud: collapse.
As long as the set of commits was not published in another repository,
it's ok (you just loose the finer steps).
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 20/04/2011 17:55, Darren New wrote:
> On 4/20/2011 9:03, Invisible wrote:
>> Doesn't appear to me that that's what happens, from what little Mercurial
>> documentation I've read.
>
> I don't know. All the mercurial documentation I've read talks about
> change sets.
The documentation I saw talks about a linear series of file versions,
just like Git and RCS and CVS and...
>> Fundamentally, VCS are about tracking changes.
>
> Fundamentally, they're about controlling versions. :-)
Well, that's a valid way to look at it I guess.
> I think it depends. If I want version 1.0 that was released, I don't
> really care what changed to get there. I want that version.
Yes, clearly.
On the other hand, if somebody sends you some stuff and says "add this
to your repo, it fixes bug #23482", you probably want to know what changed.
So yes, it does depend.
> The advantage of storing what's actually there is you can write all
> kinds of better tools to tell you the differences.
You can still apply whatever tools you want to your files, no matter
which way you store them. Although I will admit, having a diff algorithm
built right into the version control software is quite nice. (Although
sometimes I wish Darcs did this better.)
> Basically, you're storing absolutes and deducing differences, rather
> than storing differences and deducing absolutes. That means when you
> want to know what changed between release candidate 3.5RC2 and the
> version 4.2 that Fred compiled over on *his* machine, you can just
> compare the two. You don't have to reconstruct anything first.
If you just write "darcs diff", you can see the changes between any two
versions of your repo. The fact that Darcs has to do lots of work behind
the scenes to do this is of little consequence to me. Darcs has to apply
an algorithm that generates the two versions and then diffs them. Git
would have to apply an algorithm that unpacks the two commits and diffs
them. I don't really care, so long as I get my answers.
>> OK, wow. I thought having to tell Darcs when I rename stuff was
>> inconvenient, but this just sounds insane...
>
> Why? You don't have to tell git you renamed something.
Which means that it tries to guess when you rename something, so it is
100% guaranteed to guess wrong sometimes.
Still, I suppose if it's sufficiently rare, it doesn't matter too much...
>> It sounds simple enough. If this change affects line X and that change
>> affects line Y, they are independent.
>
> Yeah, until you get binary objects in there. :-)
Yeah, it's unclear how you can hope to version control a binary file,
other than just keeping a linear sequence of versions (which is what
Darcs apparently does). Personally I've never needed to try, but I guess
somebody I might.
>> And if you merge the comments branch into the main branch, and then
>> somebody
>> adds more stuff to the comments branch, then what?
>
> Then you get a merge and then more changes on the comments branch. And
> if you merge the comments branch *again*, *that* is when git uses the
> parent pointers in the commit objects to figure out which files to diff
> in order to get the patches to the parent.
>
> A--B--C--D--E--F--G--H--I
> \ | /
> Q--R/--S--T/
>
> So you started with A, changed to B, branched B and made a change to
> create Q, then R. In the mean time, I changed B to be C. Now I merge
> your R back to my C. This looks back, sees B is the common ancestor, so
> diffs R against B and applies it to C, then creates D with C and R as
> parent commits. (Each letter is a commit, which includes the state of
> the entire repository.)
>
> Now you keep working on R without incorporating my B->C change, creating
> S and T. I change D to include E and F. Now I merge your work again.
>
> Git looks at F, follows it back to D, to C and R, and sees that R is a
> common ancestor of both F and T. So it diffs T against R, applies those
> diffs to F, and creates G. You can then delete the branch that points to
> T safely without losing anything.
>
> It's *super* straightforward to understand what merges do in git.
I don't know, man, that all looks very, very complicated to me.
If I want to fix a bug in (say) GHC [which uses Darcs], I find the files
in question, edit them, record the changes, and email the file to the
GHC developers. I don't need to care about branches or whether the
development tree has changed since I got my copy of it. They don't need
to care whether my repo is in sync with theirs. They just apply the
change, and it's done. Simple.
> And if someone comes up with a better diff algorithm, no problem. The
> algorithm to do the diff during a merge isn't built into the repository.
This is only an issue for Darcs. I don't have to care how Darcs stores
my stuff. I can apply any diff algorithm I want to my files.
>> Git doesn't support recording half a file modification,
>
> Yes it does. Indeed, you can even go back and retroactively say "oh,
> those two commits? The second one should have come first, and the first
> one should be broken up into these three commits."
>
> As I said, I do this all the time.
I don't see how that's possible.
>> and doesn't even figure out which files changed.
>
> Yes it does.
Then why do you have to manually tell it which files to commit?
> It's just a two-step process. You can build up the thing you want to
> commit, and then finally commit it. It sounds like Darcs needs you to do
> that all in one step.
>
> Git does it the other way around. First it asks you what modified lines
> you want to put in the commit (and puts them in the staging area), then
> it creates the commit (based on the staging area).
I'm not sure I'm understanding what Git does. What Darcs does is show
you each change and say "do you want to put this into the commit?" If
you say yes, it records that change. If you say no, the change stays as
"new". My usual workflow when I edit stuff is to periodically run Darcs,
gather up all the changes related to one thing into a commit, run Darcs
again, gather up all the changes related to another thing into another
commit, and so on. I'm not sure what you mean by "Darcs needs you to do
that all in one step".
>>> This is trivial with GIT. I do it all the time. I'll be adding a new
>>> function, and while testing, realize there's a bug in some other
>>> function. So when everything works again, I'll do two commits, staging
>>> just particular hunks (in the diff sense of the word) and do two
>>> commits, one for the bugfix and one for the new change.
>>
>> Given that Git can only record the new file or the old one, how is that
>> possible?
>
> The staging area lies between the repository and the working directory.
So, wait, there's a third file storage area?
> So I check out some branch, and that copies it to the WD and maybe
> clears the staging area. The staging area is basically a commit that's
> not yet in the repository.
>
> Now I make changes to the WD.
>
> Then I use something like "git add" to add all the changes from the WD
> to the staging directory. Or I use "git add -i" (or, more likely, the
> GUI) to diff the WD against the staging area (or the repository), pick
> (say) three of the five diff hunks, and then create a new temp file that
> holds the repository with those three diff hunks applied, which I then
> put in the staging area. When I have everything the way I like, I commit
> the change, which copies the staging area into the repository and then
> adds a commit object pointing to it.
Damn that sounds complicated.
>> This boggles my mind. Apparently I /don't/ understand how Git works at
>> all, because the way it seems to work precludes two people touching the
>> same file at the same time...
>
> Sure. But you're thinking git tracks diffs. That's exactly the point.
I know Git doesn't track diffs - I just can't comprehend how that can
actually work properly.
> If
> I change the file, and you change the file, then now there's three
> files. The original, the new one I have, and the new one you have. When
> we go to merge it, we create number four, which is your new one with the
> differences between my version and the original applied.
This just seems a very strange way to look at things. Generally you
don't care about versions, you care about alterations. "Does this draft
have the corrections to chapter 4 in it or not?"
It seems to me that with the Git model, any time anybody edits any file,
you create a new version of the entire repo that then has to be
laboriously merged back into everybody else's repos. (Assuming no other
edits have happened in the meantime.) What a clunky way to work.
>> And the "minor detail" that if 200 people edit the same file, that's 200
>> separate branches which have to be manually merged back together again.
>
> And this differs from any other VCS how?
With a centralised system, usually it's a check-in / check-out model, so
only one person can edit a file at once.
With something like Darcs, there are now 200 change-sets, each of which
is only in some repos. Copy the change-sets around and everything is in
sync again. No need for complex "merge" operations or tangled file
histories.
> Note that if you're trying to *push* changes to a remote repository, you
> have to do it to a branch where nobody else has branched off since you
> did.
And what the hell are the chances of that ever happening? If every time
anybody touches any file it generates a new branch, then there's no
chance of ever being able to push changes back.
> In other words, if I say "update my repository to the DEV branch on
> the company's central reposityro", and I make changes, and someone else
> changes the DEV branch to point to a later version, I can no longer push
> my changes into the DEV branch. Instead, I have to fetch down the new
> DEV branch, merge my changes, then push the newly merged commit back up.
And hope that the DEV branch doesn't change while you're busy trying to
catch up. Still, I suppose if you repeat this cycle enough times,
eventually you might get lucky and be able to perform the push.
>>> If you want to merge someone's repository into yours, you simply copy
>>> from them any files or names that they have that you don't, and you're
>>> done. You're merged.
>>
>> It would be nice if Darcs worked that way.
>
> Right. In Darcs, you have to merge all the changes. In git, you have to
> merge all the changes.
No, I meant it would be nice if the Darcs repo format allowed you to
update a repo just by copying some files. Unfortunately there's
cross-references and stuff which also have to be updated, so it's not
that simple. You actually have to run Darcs to import a new patch.
Darcs also doesn't explicitly support the "bare" format that Git does,
despite it being obviously useful.
>>> Now if you want to incorporate their changes into
>>> your work, you generate a diff between their latest version and some
>>> earlier version, and apply that diff to your latest version, and you're
>>> merged.
>>
>> What a backwards way to look at it.
>
> Only if you're used to looking at source control as a series of diffs to
> start with. But that's (A) exactly what makes git hard to understand and
> (B) exactly what makes git brilliant. :-)
So doing things the hard way is brilliant?
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
>> Given that it's the repository format used by Linux developers, I think
>> it's safe to say it works adequately for multiple people editing the
>> file in parallel.
>
> This boggles my mind. Apparently I /don't/ understand how Git works at
> all...
Perhaps I can summarise:
The Darcs workflow. I download the source code, make a small edit to it,
ask Darcs to record that, and send the changes to the developers. They
add it to the central repo, which checks whether the bit I just edited
has changed since I got my copy. If not [which is quite likely], the
change is added to the central repo. Done.
The Git workflow. I download the source code, make a small edit to it,
and ask Git to record that. Git takes a complete record of every file in
the entire repo. I send that to the developers, and they try to add it
to the central repo, which makes Git check whether any unrelated changes
have happened anywhere in the entire repo since I made my change. Since
it is 100% guaranteed that this will have happened, the merge fails. I
now have to download the latest version of the source code, do a bunch
of work on my end to incorporate my 3-line edit into the newly updated
source tree, record a completely new commit object, and send that plus
the previous one back to the developers. They try to merge, find the
exact same problem, and I have to repeat all the above steps. We repeat
this endlessly until, by some fluke, I manage to execute the entire
download / merge / commit / send / have the developers merge cycle
without anybody else successfully merging to the central repo in
between. When this happens, instead of the 3-line change being added to
the central repo, we get a 25-mile string of merge commits plus the
actual 3-line edit at the beginning.
FTW, people?
Obviously this model cannot possibly work, so there must be something
I'm misunderstanding about how Git works.
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
Darren New <dne### [at] san rr com> wrote:
> Having spoken with half a dozen people who said "I hate git, it's so
> confusing" and then after showing them this they go "wow, that's really
> easy", I figured it might be worthwhile to show people this. :-)
I hate SVN. SVN projects are so easy to break accidentally, and if you
don't know the exact reason, you could be fighting to fix it for a long
time.
SVN projects are extremely fragile. SVN has the totally braindead idea
of putting a .svn project directory on each single subdirectory in the
project. (This is very unlike git, which keeps one single project directory
under the main directory where the project resides.)
For example, duplicate a directory with your favorite file manager.
Oops. SVN doesn't like that new directory at all. It refuses to add it
to the project, or do anything at all with it. If you don't know why this
happens, you are stuck. SVN refuses to do anything with it. (Solution:
Remove the .svn subdirectories from the entire offending directory hierarchy.
Most graphical file managers have no support for doing this recursively, of
course, and it's aggravated by . files being hidden in unix systems.)
Copy the contents of a directory (and its possible subdirectories) from
somewhere else (eg. update the contents of a third-party library, or copy
the work you have been doing on another platform to the project). Oops,
you just broke SVN once again. SVN will once again refuse to do anything
with this directory. You can't commit, and even an update won't fix the
problem. (Solution: Remove the entire offending directory structure, then
update, then copy the individual files from the other directory, rather
than the entire directory structure.)
Sometimes renaming/moving things from an SVN client itself can break
things, even though it shouldn't.
On a Mac the file system adds an additional layer of annoyance. Try
changing just the case of a file name (for example change "settings.hh"
to "Settings.hh") and try to figure out how to make SVN work after that.
It can be a pretty fun evening. (Not.)
--
- Warp
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 1:17, Invisible wrote:
> If you don't want
> changes you've made, you either record commits reverting them, or you just
> delete them from the history outright.
OK. It wasn't obvious from the bits I read that it was easy to delete
changes from the repository. I guess Darcs probably does that better than
mercurial or something.
> Well, yes, if you delete all your work, you have a problem. This isn't
> unique to Darcs. I'm not seeing what your point is...
That the solution to deleting branches you're in the middle of working on is
the same in both cases: "Duh, don't do that. Or make backups."
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 2:26, Invisible wrote:
> You can still apply whatever tools you want to your files, no matter which
> way you store them. Although I will admit, having a diff algorithm built
> right into the version control software is quite nice. (Although sometimes I
> wish Darcs did this better.)
That's exactly what I mean. In order to have Darcs do this better, you have
to actually fix your repo to use the better algorithm. In git, the algorithm
isn't part of the repo. It's only part of the tools. Your very statement
"it's built in, but I wish Darcs did it better" is exactly my point.
> If you just write "darcs diff", you can see the changes between any two
> versions of your repo.
What if I want to use kdiff3 or gvimdiff or some other diff visualization tool?
>>> OK, wow. I thought having to tell Darcs when I rename stuff was
>>> inconvenient, but this just sounds insane...
>>
>> Why? You don't have to tell git you renamed something.
>
> Which means that it tries to guess when you rename something, so it is 100%
> guaranteed to guess wrong sometimes.
If a file disappears from one place and at the same time reappears somewhere
else with exactly to the byte the same contents, does it really matter
whether you renamed it, whether you cut and pasted, or whether you typed it
back in?
Remember, git is storing snapshots of the repository, not changes. The fact
that something got renamed is sort of irrelevant.
Darcs is also storing snapshots, except snapshots of changes. If I rename
file A to B, and then from B to C, and *then* I commit, Darcs isn't going to
have those changes either. If I delete five lines, then another ten, then
commit, Darcs is going to lose that history that that was actually two changes.
> Yeah, it's unclear how you can hope to version control a binary file, other
> than just keeping a linear sequence of versions (which is what Darcs
> apparently does). Personally I've never needed to try, but I guess somebody
> I might.
Word documents. Images. Audio. Video game resources. People want version
control for all of that.
Keeping all the old versions is the way to do it. The problem comes when you
have a distributed repo, and you have to store locally every old version of
all the binary files that you almost never are going to want.
Imagine if you were a Linux developer and people stored installation CD ISO
images in the repository. Do you really want to check out every copy of
every install CD just so you can fix bugs in one file system?
> I don't know, man, that all looks very, very complicated to me.
You asked what happens in that case. That's what happens.
> If I want to fix a bug in (say) GHC [which uses Darcs], I find the files in
> question, edit them, record the changes, and email the file to the GHC
> developers. I don't need to care about branches or whether the development
> tree has changed since I got my copy of it. They don't need to care whether
> my repo is in sync with theirs. They just apply the change, and it's done.
> Simple.
That's how it works with git also. Indeed, there's a git command that says
"generate an email with the patch in it that I need to update someone else's
repository."
You are making a branch. Your whole repository is a branch of the other
guy's repository. If you look at the up-pointing lines as "you mail me a
patch", then you get the same answer.
You asked what happens when someone keeps working on a branch that someone
else already incorporated. I showed you how git decides which diffs to apply
and which not to apply. Darcs does the same thing when building the working
directory. It's going to apply the new patches, but how does it know what
the new patches are? Right, it goes back until it finds the patches it
already applied, then applies the newer ones.
>>> Git doesn't support recording half a file modification,
>>
>> Yes it does. Indeed, you can even go back and retroactively say "oh,
>> those two commits? The second one should have come first, and the first
>> one should be broken up into these three commits."
>>
>> As I said, I do this all the time.
>
> I don't see how that's possible.
Here's how to do it without the GUI:
http://book.git-scm.com/4_interactive_adding.html
With a GUI, you look at the patch list, right-click a hunk, and say "stage
this to be committed".
If you want to change old commits, you do this:
http://book.git-scm.com/4_interactive_rebasing.html
Again, that's the text-based way of doing it without a gui.
>>> and doesn't even figure out which files changed.
>>
>> Yes it does.
>
> Then why do you have to manually tell it which files to commit?
Because maybe you don't want to commit all your changes in one step.
> I'm not sure I'm understanding what Git does. What Darcs does is show you
> each change and say "do you want to put this into the commit?" If you say
> yes, it records that change. If you say no, the change stays as "new".
In git, say you start with a working directory that matches the latest thing
in the repository. You change files AA, BB, and CCC, and you add file DD.
Changing the first two were to fix a bug, and the second two added a
configuration option.
git add AA BB
git commit -m "fix bug"
git add CCC DD
git commit -m "add configuration option"
Let's say you then change 173 files, converting all single quotes to double
quotes. You can then say
git add -a
git commit -m "change quote style"
> not sure what you mean by "Darcs needs you to do that all in one step".
I mean that gathering up the changes and committing them sounds like a
single step in Darcs. In git, I can say
wings3d my_model.wings
git add my_model.wings
gimp my_image.jpg
git add my_image.jpg
vi configuration.ini
git add configuration.ini
git commit -m "add a textured model with some configuration"
>> The staging area lies between the repository and the working directory.
> So, wait, there's a third file storage area?
Yes. That's where you build up the next commit. It's called the staging
area, or the index.
>> Then I use something like "git add" to add all the changes from the WD
>> to the staging directory. Or I use "git add -i" (or, more likely, the
>> GUI) to diff the WD against the staging area (or the repository), pick
>> (say) three of the five diff hunks, and then create a new temp file that
>> holds the repository with those three diff hunks applied, which I then
>> put in the staging area. When I have everything the way I like, I commit
>> the change, which copies the staging area into the repository and then
>> adds a commit object pointing to it.
>
> Damn that sounds complicated.
It's very simple with the gui. You start up the gui, it shows you a top-left
pane of files that have changed that you haven't decided to commit yet.
Bottom right pane are files that'll be in the commit. Right side is the
listing of the diffs for whatever file you've highlighted.
If you want to put half the changes from configuration.ini in your commit,
you click on that, go over to the list of diffs, click on each one you want
in the commit, then click the commit button.
Pretty trivial.
>>> This boggles my mind. Apparently I /don't/ understand how Git works at
>>> all, because the way it seems to work precludes two people touching the
>>> same file at the same time...
>>
>> Sure. But you're thinking git tracks diffs. That's exactly the point.
>
> I know Git doesn't track diffs - I just can't comprehend how that can
> actually work properly.
Because files are actually named by their SHA-1, so nobody ever touches two
different copies of the same file at the same time.
>> If
>> I change the file, and you change the file, then now there's three
>> files. The original, the new one I have, and the new one you have. When
>> we go to merge it, we create number four, which is your new one with the
>> differences between my version and the original applied.
>
> This just seems a very strange way to look at things. Generally you don't
> care about versions, you care about alterations. "Does this draft have the
> corrections to chapter 4 in it or not?"
And you can trivially tell that in git, not by looking at the files, but by
looking at the commits.
> It seems to me that with the Git model, any time anybody edits any file, you
> create a new version of the entire repo that then has to be laboriously
> merged back into everybody else's repos. (Assuming no other edits have
> happened in the meantime.) What a clunky way to work.
It's not laborious to merge it in, any more than it's laborious to merge
changes in Darcs into your repository and working directory.
If I clone a repository from you, your repository's URL is stored in my
repository and called "origin" (by default). If I want to fetch all your
changes, I say "git pull origin", which connects to your repository, gets
the list of objects you have that I don't, and pulls them down. It also
updates any branch names that you changed since last time I did that, so if
you have a branch called "bugfix", I'll have a branch called
"origin/bugfix". If Sally also cloned your repository and created a branch
called bugfix, then I pulled from Sally also, I'll have a branch called
"origin/bugfix" and one called "sally/bugfix".
>>> And the "minor detail" that if 200 people edit the same file, that's 200
>>> separate branches which have to be manually merged back together again.
>>
>> And this differs from any other VCS how?
>
> With a centralised system, usually it's a check-in / check-out model, so
> only one person can edit a file at once.
Um, no, not for the last 15 years or so. Not even CVS did things that way,
let alone SVN. Some systems work like that, yes, but they work really,
really poorly when you have 200 people working on the same files, which is
why people moved to CVS in the first place.
> With something like Darcs, there are now 200 change-sets, each of which is
> only in some repos. Copy the change-sets around and everything is in sync
> again. No need for complex "merge" operations or tangled file histories.
Of course you need to merge them, and of course you'll have tangled file
histories. If all 200 people change the same part of the file, you'll have
200 merge conflicts. If everyone is passing around partial change sets and
making more changes that are dependent on those changes, you'll have a
tangled file history.
>> Note that if you're trying to *push* changes to a remote repository, you
>> have to do it to a branch where nobody else has branched off since you
>> did.
>
> And what the hell are the chances of that ever happening? If every time
> anybody touches any file it generates a new branch, then there's no chance
> of ever being able to push changes back.
http://www.kernel.org/pub/software/scm/git/docs/git-rebase.html
Basically, you say "go look at the changes I made since I branched off the
upstream repository, then apply those same changes to the new head of the
upstream repository, and submit *that* as the new commit."
> And hope that the DEV branch doesn't change while you're busy trying to
> catch up. Still, I suppose if you repeat this cycle enough times, eventually
> you might get lucky and be able to perform the push.
If the DEV branch is changing that quickly, it means *someone* is going
through this cycle successfully. You're complaining that nobody goes to
that restaurant any more because it's always too crowded.
The rebase is a single step, unless there are merge conflicts, so it's
basically bound by network and CPU.
>>>> If you want to merge someone's repository into yours, you simply copy
>>>> from them any files or names that they have that you don't, and you're
>>>> done. You're merged.
>>>
>>> It would be nice if Darcs worked that way.
>>
>> Right. In Darcs, you have to merge all the changes. In git, you have to
>> merge all the changes.
>
> No, I meant it would be nice if the Darcs repo format allowed you to update
> a repo just by copying some files.
Well, git *does* store stuff in files, so technically you could copy the
files. But by "copy files" I mean "use git to copy the new files." As in,
"you don't have to run any diffs or patches or anything".
> Darcs also doesn't explicitly support the "bare" format that Git does,
> despite it being obviously useful.
Which is why you don't have trouble with the rebasing. You only need to
rebase stuff when you're pushing to a bare repository without human
intervention. Basically, if you're sending changes to a bare repository
without human intervention, you have to prove you've already resolved the
merge conflicts that such might impose on someone else who later updates
from that bare repository.
>>>> Now if you want to incorporate their changes into
>>>> your work, you generate a diff between their latest version and some
>>>> earlier version, and apply that diff to your latest version, and you're
>>>> merged.
>>>
>>> What a backwards way to look at it.
>>
>> Only if you're used to looking at source control as a series of diffs to
>> start with. But that's (A) exactly what makes git hard to understand and
>> (B) exactly what makes git brilliant. :-)
>
> So doing things the hard way is brilliant?
Building an entire type theory to discuss isomorphic idempotent change sets
is the easy way?
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 2:35, Invisible wrote:
>>> Given that it's the repository format used by Linux developers, I think
>>> it's safe to say it works adequately for multiple people editing the
>>> file in parallel.
>>
>> This boggles my mind. Apparently I /don't/ understand how Git works at
>> all...
>
> Perhaps I can summarise:
>
> The Darcs workflow. I download the source code, make a small edit to it, ask
> Darcs to record that, and send the changes to the developers. They add it to
> the central repo, which checks whether the bit I just edited has changed
> since I got my copy. If not [which is quite likely], the change is added to
> the central repo. Done.
That's exactly how git works, workflow-wise, if you want.
http://book.git-scm.com/5_git_and_email.html
You *also* have the option of telling the other guy the URL of your
repository and having him suck it in, or an option of pushing to a (possibly
bare and unattended) repository your changes.
> The Git workflow. I download the source code, make a small edit to it, and
> ask Git to record that. Git takes a complete record of every file in the
> entire repo.
Uh, no. Not if you're doing it manually. If you're automating it by cloning
the repository or pulling and pushing changes over git protocols without
human intervention, then yes. But since your repository tracks what you got
from other people, you already know about changes you made that they don't have.
> I send that to the developers, and they try to add it to the
> central repo, which makes Git check whether any unrelated changes have
> happened anywhere in the entire repo since I made my change. Since it is
> 100% guaranteed that this will have happened, the merge fails.
Wow. Advice: Don't talk to anyone about how git works, since even if you
think you've figured it out, you haven't.
> Obviously this model cannot possibly work, so there must be something I'm
> misunderstanding about how Git works.
Ah, thank you.
Yes. You're thinking that you can't add a commit to a repository without
merging it in. You're still thinking there's one authoritative set of files.
There aren't. There's bunches of snapshots.
So when you send your change to the maintainers, you're sending them new
files. There are no merge conflicts, because there is no merge. Every
changed file has a new sha-1, so it's a new file. git can look at the
repository and know where in your working directory each file belongs
(because some of the files in the repository are directories), but in the
repository itself it's all just a big flat bag of sha-1 files.
So sending your changes to someone else doesn't cause merge conflicts, and
there's no need to back them out.
When I want to put your changes into *my* version of the files, I have to
merge it, just like Darcs.
The equivalent description would be that if you changed something in Darcs
nd I made a change in my working directory, you could no longer send me
patches until I threw away all my working-directory changes, in case there
was a conflict.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 7:46, Warp wrote:
> I hate SVN. SVN projects are so easy to break accidentally, and if you
> don't know the exact reason, you could be fighting to fix it for a long
> time.
I agree. I've *never* gotten an svn branch to merge back. I have *always*
built a patch, and then applied the patch to the head.
> SVN projects are extremely fragile. SVN has the totally braindead idea
> of putting a .svn project directory on each single subdirectory in the
> project. (This is very unlike git, which keeps one single project directory
> under the main directory where the project resides.)
While it lets you check out only parts of the project, it's a PITA to deal
with. Especially when you want to recursively grep your code for the last
vestiges of variable abcxyz, and it's all over inside the files under .svn.
git doesn't do that because a git repository is basically a big flat bag of
files with no possibility of conflicting file names, so there's no
subdirectories possible. :-)
But yah, all that stuff is a pain. I wound up having to toss and check out
the entire local copy at least once a month at work, just because crap would
break and I couldn't fix it locally.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 2:35, Invisible wrote:
> The Darcs workflow. I download the source code, make a small edit to it, ask
> Darcs to record that, and send the changes to the developers. They add it to
> the central repo, which checks whether the bit I just edited has changed
> since I got my copy. If not [which is quite likely], the change is added to
> the central repo. Done.
OK. All the confusions you think you're seeing in git are due to Darcs just
not being able to do what git does.
When Darcs gets a merge conflict, it just doesn't apply *either* patch.
"Darcs escapes this problem by ignoring those parts of the patches that
conflict."
Obviously, if Darcs isn't going to try to fix the patches for you, it's a
lot easier to record a bunch of conflicting patches. The equivalent in git
is to pull in changes from other repositories and then not trying to update
your working directory to include the changes. Trivial one-liner that always
works.
"If the conflict is with one of your not-yet-published patches, you may
choose to amend that patch rather than creating a resolve patch."
And that's exactly what the "git merge" command does. It takes your patches,
and someone else's patches, and merges them together. It's a separate step
because, unlike Darcs, git supports having more than one "pristine" in the
same repository. If you only have one branch in git, it's as trivial as only
having one repository in Darcs.
"This is how a project with many contributors, but every contribution is
reviewed and manually applied by the project leader, can be run." This is
the bit about sending email you were talking about. git can work that way,
and the terrible "merge" problems you're talking about are handled the same
way: the guy getting the patches fixes the merge.
What Darcs apparently can't do is support any way of doing distributed
development with an authoritative repository *without* someone dedicated to
fixing the merge conflicts. That's where the whole rant you're talking about
came from. If you have someone who is going to look at your changes (i.e.,
if I am CTO and I want to build a new release or something) then the whole
"fast-forward merge with a rebase" dance is unnecessary and indeed
counterproductive, as it loses the history of the change (in the same sense
that unrecording a bunch of changes and re-recording one big change in Darcs
loses the history).
I'll grant you that Darcs is definitely simpler, but I think it's less
capable also, and that's the primary place the simplicity comes from. The
Darcs replace command is interesting, but I'm not sure how well that would
work in practice, especially in languages with complex scoping.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
>> If you just write "darcs diff", you can see the changes between any two
>> versions of your repo.
>
> What if I want to use kdiff3 or gvimdiff or some other diff
> visualization tool?
Then yeah, you'd have to construct two versions of the repo, just like
you would with any other version control system.
>> Yeah, it's unclear how you can hope to version control a binary file,
>> other
>> than just keeping a linear sequence of versions (which is what Darcs
>> apparently does). Personally I've never needed to try, but I guess
>> somebody
>> I might.
>
> Word documents. Images. Audio. Video game resources. People want version
> control for all of that.
Trying to use version control for Word documents would be a fairly
insane thing to do. (Unless you can actually parse the binary file and
ignore all the metadata...) But sure, I can see how if I was doing
something other than writing text, I'd probably want to version control
that too.
> Imagine if you were a Linux developer and people stored installation CD
> ISO images in the repository. Do you really want to check out every copy
> of every install CD just so you can fix bugs in one file system?
Man, version controlling a multi-GB ISO image sounds like a barrel of
laughs. o_O
>>>> and doesn't even figure out which files changed.
>>>
>>> Yes it does.
>>
>> Then why do you have to manually tell it which files to commit?
>
> Because maybe you don't want to commit all your changes in one step.
So you have to manually remember what you changed? That seems rather...
primitive. I much prefer the way Darcs does it (i.e., prompt you for
which changes you want to include in this particular commit).
>> not sure what you mean by "Darcs needs you to do that all in one step".
>
> I mean that gathering up the changes and committing them sounds like a
> single step in Darcs.
A single interactive step, yes.
>> Damn that sounds complicated.
>
> It's very simple with the gui. You start up the gui, it shows you a
> top-left pane of files that have changed that you haven't decided to
> commit yet. Bottom right pane are files that'll be in the commit. Right
> side is the listing of the diffs for whatever file you've highlighted.
>
> If you want to put half the changes from configuration.ini in your
> commit, you click on that, go over to the list of diffs, click on each
> one you want in the commit, then click the commit button.
>
> Pretty trivial.
So Git has a GUI tool that lets you do what Darcs does natively?
>> With a centralised system, usually it's a check-in / check-out model, so
>> only one person can edit a file at once.
>
> Some systems work like that, yes, but they work
> really, really poorly when you have 200 people working on the same
> files
Presumably this is why everybody wants distributed version control now.
>> With something like Darcs, there are now 200 change-sets, each of
>> which is
>> only in some repos. Copy the change-sets around and everything is in sync
>> again. No need for complex "merge" operations or tangled file histories.
>
> Of course you need to merge them, and of course you'll have tangled file
> histories. If all 200 people change the same part of the file, you'll
> have 200 merge conflicts. If everyone is passing around partial change
> sets and making more changes that are dependent on those changes, you'll
> have a tangled file history.
Except that usually 200 people will be editing 200 different parts of
the repository.
> Basically, you say "go look at the changes I made since I branched off
> the upstream repository, then apply those same changes to the new head
> of the upstream repository, and submit *that* as the new commit."
The Darcs model is "the bugfix for #5326 is *this* alteration". You can
then apply that to whatever you want. (Unless the modified bit is
altered, of course.) The Git model seems to be to do lots of extra work
and then record it as a new item of data, which you don't actually need,
but that's just hot Git works.
>> No, I meant it would be nice if the Darcs repo format allowed you to
>> update a repo just by copying some files.
>
> Well, git *does* store stuff in files, so technically you could copy the
> files. But by "copy files" I mean "use git to copy the new files." As
> in, "you don't have to run any diffs or patches or anything".
Oh, I see.
>> So doing things the hard way is brilliant?
>
> Building an entire type theory to discuss isomorphic idempotent change
> sets is the easy way?
After you've built it, yes. The theory is hard, but it makes using the
tool easy.
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
> Wow. Advice: Don't talk to anyone about how git works, since even if you
> think you've figured it out, you haven't.
>
>> Obviously this model cannot possibly work, so there must be something I'm
>> misunderstanding about how Git works.
>
> Ah, thank you.
Like I said, the way Git appears to work is utterly hopeless, so I must
have got something wrong somewhere.
> So when you send your change to the maintainers, you're sending them new
> files. There are no merge conflicts, because there is no merge.
Right. Every time you edit a file, it's actually a new file.
> So sending your changes to someone else doesn't cause merge conflicts,
> and there's no need to back them out.
OK.
> When I want to put your changes into *my* version of the files, I have
> to merge it, just like Darcs.
1. Having a copy of someone else's changes is useless unless I can
incorporate them into the latest version of the files. This is the
equivilent of three people taking copies of a file, editing it, and
emailing me back three modified files. I don't want three files, I want
*one* which has all the edits in it.
2. If I do "darcs pull", all new history is copied to my repo, and the
working copy is updated to reflect these changes. [Assuming there are no
conflicts of course.] That's it. That's all you have to do.
With Git, you have to do some crazy thing with comparing my working copy
to the most recent common ancestor of my branch and the branch I'm
pulling, creating diffs for that, comparing the most recent common
ancestor to the snapshot I just pulled, creating a diff for that,
combining the two diffs, and then applying that to the most recent
common ancestor. You then have to *create a new commit object*
representing this new combined state.
Every time you try to combine two states of the repo, it creates a new
commit object representing the merge. Darcs, by contrast, lets me
trivially apply any combination of changes I want. I can even create
files with combinations of changes that have never existed before if I like.
I guess the thing that really flips my lid is that not only does Git
require you to construct a useless merge commit every time you want to
do something as trivial as put two changes together, but if you want to
put new stuff into a repo, you have to somehow get the merge up to date
first. That sounds increadibly fragile.
It also irritates me that Git insists that even unrelated changes must
have a linear time ordering. Still, it's not like I actually have to
*use* Git, so it doesn't really matter if I don't like it...
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 13:07, Orchid XP v8 wrote:
> Then yeah, you'd have to construct two versions of the repo, just like you
> would with any other version control system.
Except git. Because git doesn't store diffs, it stores files. :-) So it's
trivial for git to run whatever you want against two different versions of
the files.
> Man, version controlling a multi-GB ISO image sounds like a barrel of
> laughs. o_O
That's one of the reasons why people like centralized servers.
>>>>> and doesn't even figure out which files changed.
>>>>
>>>> Yes it does.
>>>
>>> Then why do you have to manually tell it which files to commit?
>>
>> Because maybe you don't want to commit all your changes in one step.
>
> So you have to manually remember what you changed?
No, of course not. Are you even reading what I'm writing? Why would the
fact that you tell git what changes you want to commit with one command then
actually commit those with a second command mean that git can't tell what
you changed?
> I much prefer the way Darcs does it (i.e., prompt you for which
> changes you want to include in this particular commit).
I prefer the GUI, actually. Much easier to pick out various commits to commit.
>>> not sure what you mean by "Darcs needs you to do that all in one step".
>>
>> I mean that gathering up the changes and committing them sounds like a
>> single step in Darcs.
>
> A single interactive step, yes.
Which means you can't (for example) stop in the middle when you realize you
forgot to make one of the 30 changes you want to commit to fix something in
particular. You have to start over.
> So Git has a GUI tool that lets you do what Darcs does natively?
It's native to git too. You get to use the command line or a gui. You're
really not actually reading what I'm writing, are you?
> Except that usually 200 people will be editing 200 different parts of the
> repository.
And if that's the case in git, then you have no trouble merging things when
and as you want them.
>> Basically, you say "go look at the changes I made since I branched off
>> the upstream repository, then apply those same changes to the new head
>> of the upstream repository, and submit *that* as the new commit."
>
> The Darcs model is "the bugfix for #5326 is *this* alteration". You can then
> apply that to whatever you want. (Unless the modified bit is altered, of
> course.) The Git model seems to be
The git model seems to be something you're still fully unfamiliar with.
> to do lots of extra work and then record
> it as a new item of data, which you don't actually need, but that's just hot
> Git works.
What extra work? If you just want to commit everything you've changed, you
say "git commit -all" or some such, and away you go. If you want to take my
changes and update your repository, you say "git pull darren", and when
you're ready to incorporate my changes into your development, you say "git
merge darren". It's two steps because you don't want to tie "get Darren's
changes" to "make sure Darren's changes are all compatible with mine."
>>> So doing things the hard way is brilliant?
>>
>> Building an entire type theory to discuss isomorphic idempotent change
>> sets is the easy way?
>
> After you've built it, yes. The theory is hard, but it makes using the tool
> easy.
Honestly, I'm not sure I see any advantage of Darcs over git, other than
possibly repository size. I imagine once you get enough change sets in a
repository (think 10 years of Linux development), it could get really slow
to check out a particular file.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 13:28, Orchid XP v8 wrote:
> Right. Every time you edit a file, it's actually a new file.
In the repository, yes. Obviously, git knows that there's a file in your
repository with the same path and name as a file in my repository but
different contents.
>> When I want to put your changes into *my* version of the files, I have
>> to merge it, just like Darcs.
>
> 1. Having a copy of someone else's changes is useless unless I can
> incorporate them into the latest version of the files.
Which latest versions? Oh, right, Darcs only has one latest version, and if
you don't want to apply changes from someone, you can't fetch them either.
> This is the
> equivilent of three people taking copies of a file, editing it, and emailing
> me back three modified files. I don't want three files, I want *one* which
> has all the edits in it.
That's the merge step. Three people send you three copies. When you're
ready, you say "apply those changes to my copy." I'm not sure where your
confusion is. In Darcs, that's one step it seems - I can't get changes from
you *without* applying them to the work I'm doing. In git, it's two steps,
because maybe you're in the middle of something and you don't want to merge
in my changes until the stuff you're working on actually works and passes
tests and stuff.
> 2. If I do "darcs pull", all new history is copied to my repo, and the
> working copy is updated to reflect these changes. [Assuming there are no
> conflicts of course.] That's it. That's all you have to do.
So if you have something like Linux, where there's a new release every few
months, you need a complete repository for every release. And if I fix a bug
in an old release and you want to incorporate that bug fix into newer
releases, what do you do?
> With Git, you have to do some crazy thing with comparing my working copy to
> the most recent common ancestor of my branch and the branch I'm pulling,
> creating diffs for that, comparing the most recent common ancestor to the
> snapshot I just pulled, creating a diff for that, combining the two diffs,
> and then applying that to the most recent common ancestor.
git does all that. Complaining about that is like me complaining "To check
out a working copy from Darcs, I have to do some crazy thing about
topological sorts on patches, then applying them to empty files in the right
order, it's nuts!"
> You then have to
> *create a new commit object* representing this new combined state.
Yes.
> Every time you try to combine two states of the repo, it creates a new
> commit object representing the merge. Darcs, by contrast, lets me trivially
> apply any combination of changes I want. I can even create files with
> combinations of changes that have never existed before if I like.
You can do that in git too. You have to create a commit if you want to store
it in the repository, tho.
> I guess the thing that really flips my lid is that not only does Git require
> you to construct a useless merge commit every time you want to do something
> as trivial as put two changes together,
Commits are trivially inexpensive in git. They represent the state of some
particular repository at some particular time. It's how you store stuff.
Say you have 25 changes in your Darcs repository, and you want a version
that applies every change except #23. What do you put in the repository to
represent that?
> but if you want to put new stuff
> into a repo, you have to somehow get the merge up to date first.
No. You really aren't listening to what I'm saying, so I'm not sure why I'm
bothering.
You only have to do that step *IF* nobody is going to look for merge
conflicts. I.e., *you* have to do that step *only* in the case that *you*
want to change *my* repository without *me* being there. Which is something
Darcs can't do at all.
> It also irritates me that Git insists that even unrelated changes must have
> a linear time ordering.
No they don't. Indeed, if unrelated changes had to have a linear time
ordering, you wouldn't get merge conflicts at all, would you? Then you
wouldn't be complaining about git having merge commits.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 14:46, Darren New wrote:
>> So Git has a GUI tool that lets you do what Darcs does natively?
>
> It's native to git too.
Well, actually, the thing is, git has a bunch of layers. There's the layer
to just put a file into the staging index, a command to create a commit from
the staging index, a command to point a particular name at a particular
commit object, etc.
Everything beyond that, including deciding what hunks get included in those
files and so on, is at the next layer up. In exactly the same way that
Darcs runs on top of the file system.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 13:28, Orchid XP v8 wrote:
> 2. If I do "darcs pull", all new history is copied to my repo, and the
> working copy is updated to reflect these changes. [Assuming there are no
> conflicts of course.] That's it. That's all you have to do.
OK, so say I have a Darcs repository and I'm working on my program. I have
some changes in my working directory, when I come across a bug. I ask you
about it, and you say "that's already fixed." Can I pull just that one
change into my repository?
What if I'm working on version 1 of the program, and you made a fix in
version 2? How do I handle that?
Indeed, how do I get out a version of the working directory excluding just
one particular patch back in time, in order to see (for example) if that
patch was the cause of a bug? I didn't see that in the Darcs manual. All I
saw was deleting changes from the repository.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 13:07, Orchid XP v8 wrote:
>> Well, git *does* store stuff in files, so technically you could copy the
>> files. But by "copy files" I mean "use git to copy the new files." As
>> in, "you don't have to run any diffs or patches or anything".
>
> Oh, I see.
Actually, you can store a repository on an http server and give everyone
read access to the repository, and everyone can update from there. So, yeah,
it basically *can* be done by copying the right files from one place to another.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 13:28, Orchid XP v8 wrote:
> 1. Having a copy of someone else's changes is useless unless I can
> incorporate them into the latest version of the files.
Also, does Darcs record when you applied various patches?
So if, for example, I'm working, and everything's good, and I take some
patches from you, then work some more, then take some patches from Sam, then
work some more, then run my test and it fails, can I figure out that it was
Sam's patches, even if he created those patches before I even cloned the
repository in the first place?
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/21/2011 13:28, Orchid XP v8 wrote:
> Every time you try to combine two states of the repo, it creates a new
> commit object representing the merge.
Just to be clear, adding your changes to my repository is just adding a
commit. A commit is like a Darcs patch. One update of the repository is one
commit.
There's a "merge commit" which is nothing but a commit saying "this is what
it looks like after you merge two commits."
That said, the Linux repository going back 234 tagged versions (back to
2.6.11, which isn't all that far back) has 244,000 commits, of which only
15,000 are merges. So people tend to make 10 or more changes on each branch
before they merge it into the repository.
I'm not sure I'd want to check out something from Darcs that has a quarter
million patches in it and wait for Darcs to apply them all one by one. How
well does it handle that?
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 22/04/2011 12:39 AM, Darren New wrote:
> On 4/21/2011 13:07, Orchid XP v8 wrote:
>>> Well, git *does* store stuff in files, so technically you could copy the
>>> files. But by "copy files" I mean "use git to copy the new files." As
>>> in, "you don't have to run any diffs or patches or anything".
>>
>> Oh, I see.
>
> Actually, you can store a repository on an http server and give everyone
> read access to the repository, and everyone can update from there. So,
> yeah, it basically *can* be done by copying the right files from one
> place to another.
I've done this with Darcs. Unfortunately, to update the repo, you
basically delete it off the server and copy the current version over
there - which is tedious. (By default, the server would also have a
working copy - which is pointless. Fortunately, if you delete that bit,
it still works.)
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
>> I much prefer the way Darcs does it (i.e., prompt you for which
>> changes you want to include in this particular commit).
>
> I prefer the GUI, actually. Much easier to pick out various commits to
> commit.
I will admit, a GUI is probably a superior interface for doing this.
(And I'm not aware of anyone having built a GUI for Darcs.)
>>>> not sure what you mean by "Darcs needs you to do that all in one step".
>>>
>>> I mean that gathering up the changes and committing them sounds like a
>>> single step in Darcs.
>>
>> A single interactive step, yes.
>
> Which means you can't (for example) stop in the middle when you realize
> you forgot to make one of the 30 changes you want to commit to fix
> something in particular. You have to start over.
True enough. It would be nice if you could make Darcs remember which
changes you selected last time around. (Then again, some of those
changes might no longer exist next time you run Darcs...)
>> Except that usually 200 people will be editing 200 different parts of the
>> repository.
>
> And if that's the case in git, then you have no trouble merging things
> when and as you want them.
You still have to manually make Git combine all 200 changes into one
version, and then commit that. It just seems like an undless cycle of
merges trying to keep everything straight, generating an ever more
tangled history behind it.
>> to do lots of extra work and then record
>> it as a new item of data, which you don't actually need, but that's
>> just hot
>> Git works.
>
> What extra work? If you just want to commit everything you've changed,
> you say "git commit -all" or some such, and away you go. If you want to
> take my changes and update your repository, you say "git pull darren",
> and when you're ready to incorporate my changes into your development,
> you say "git merge darren". It's two steps because you don't want to tie
> "get Darren's changes" to "make sure Darren's changes are all compatible
> with mine."
It's more the conceptual annoyance of having to record every merge
operation as a new version of the entire repo, even if you only changed
one line. That seems really clumsy to me.
> Honestly, I'm not sure I see any advantage of Darcs over git.
Likewise, but inverted.
Still, until GHC moves from Darcs to Git, I won't have to actually care,
so I guess it doesn't really matter.
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 22/04/2011 12:31 AM, Darren New wrote:
> OK, so say I have a Darcs repository and I'm working on my program. I
> have some changes in my working directory, when I come across a bug. I
> ask you about it, and you say "that's already fixed." Can I pull just
> that one change into my repository?
Obviously yes.
> What if I'm working on version 1 of the program, and you made a fix in
> version 2? How do I handle that?
Depends which way I set it up. If each version of the program is just a
tag, then see the above. If each version is a seperate repo, you just
need to pull from the v2 repo. (How much conflict you'll have to deal
with is another matter - but that's the fun of backporting, eh?)
> Indeed, how do I get out a version of the working directory excluding
> just one particular patch back in time, in order to see (for example) if
> that patch was the cause of a bug?
You say to Darcs "please revert patch #34823". If that doesn't fix your
problem, you say to Darcs "please revert all unrecorded changes" (since
the inverse change you just asked for hasn't been recorded yet). If that
*does* fix your problem... well, you can just hit record, or you can
investigate further to see if you can keep the good parts of the
original patch without causeing problems, or whatever.
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 22/04/2011 12:46 AM, Darren New wrote:
> On 4/21/2011 13:28, Orchid XP v8 wrote:
>> 1. Having a copy of someone else's changes is useless unless I can
>> incorporate them into the latest version of the files.
>
> Also, does Darcs record when you applied various patches?
Not to my knowledge, no. It records who created a patch and when, but
not when it was applied to any particular repo.
> So if, for example, I'm working, and everything's good, and I take some
> patches from you, then work some more, then take some patches from Sam,
> then work some more, then run my test and it fails, can I figure out
> that it was Sam's patches, even if he created those patches before I
> even cloned the repository in the first place?
In that case you're presumably going to revert patches until the problem
goes away. Maybe one patch broke something, maybe its an interaction of
several patches. You turn patches on and off until you figure out what's up.
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 22/04/2011 11:29 PM, Darren New wrote:
> On 4/21/2011 13:28, Orchid XP v8 wrote:
>> Every time you try to combine two states of the repo, it creates a new
>> commit object representing the merge.
>
> Just to be clear, adding your changes to my repository is just adding a
> commit. A commit is like a Darcs patch. One update of the repository is
> one commit.
>
> There's a "merge commit" which is nothing but a commit saying "this is
> what it looks like after you merge two commits."
Seems more like because it's so difficult to merge two files, after
you've done it you have to save it to prevent you having to redo all
that complex hard work. And in the process, all change application is
forced to become strictly linear.
Apparently this doesn't stop people working on the Linux kernel. But it
seems really clumsy to me.
> I'm not sure I'd want to check out something from Darcs that has a
> quarter million patches in it and wait for Darcs to apply them all one
> by one. How well does it handle that?
You're aware that Darcs keeps a cached copy of the latest state of all
the files, so it doesn't have to recompute them, right?
Last time I tried downloading the repos for GHC, it was dominated by
network latency. Processor usage was almost non-existent. It just takes
a long time to shift gigabytes of data over a slow ADSL link. Just as it
would if I had downloaded a Zip file of the source code with no history
data at all.
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
>> 1. Having a copy of someone else's changes is useless unless I can
>> incorporate them into the latest version of the files.
>
> Which latest versions? Oh, right, Darcs only has one latest version, and
> if you don't want to apply changes from someone, you can't fetch them
> either.
Hint: Look up "darcs fetch". It fetches changes without applying them.
> That's the merge step. Three people send you three copies. When you're
> ready, you say "apply those changes to my copy." I'm not sure where your
> confusion is.
It's the fact that you have to commit the merged version of the file as
a new version. Every time you merge in a new change, you have to commit
a new version. It just seems clunky and unecessary.
> In Darcs, that's one step it seems - I can't get changes
> from you *without* applying them to the work I'm doing.
Not true. You can get changes without applying them. It's just that
usually, you want to fetch changes to, you know, *use* them.
> In git, it's two
> steps, because maybe you're in the middle of something and you don't
> want to merge in my changes until the stuff you're working on actually
> works and passes tests and stuff.
Then why fetch the changes at all? Why not wait until you're actually
ready to apply them?
> So if you have something like Linux, where there's a new release every
> few months, you need a complete repository for every release.
Um... why?
> And if I
> fix a bug in an old release and you want to incorporate that bug fix
> into newer releases, what do you do?
Oh, I see. You mean if you actually have multiple versions of something
being developed concurrently? Yeah, in that case you'd have to move the
changeset from one branch to another and hope it works.
>> You then have to
>> *create a new commit object* representing this new combined state.
>
> Yes.
So if two people have the same repo, and they merge in change X and then
later merge in change Y, and then the other guy merges in change Y first
and later change X, they now apparently have conflicting histories. (And
all because Git wants to pretend that everything happens in linear
order.) How do you get out of that?
>> I guess the thing that really flips my lid is that not only does Git
>> require
>> you to construct a useless merge commit every time you want to do
>> something
>> as trivial as put two changes together,
>
> Commits are trivially inexpensive in git. They represent the state of
> some particular repository at some particular time. It's how you store
> stuff.
As I say, it just annoys me that you have to assign an arbitrary
ordering to changes.
> Say you have 25 changes in your Darcs repository, and you want a version
> that applies every change except #23. What do you put in the repository
> to represent that?
You ask Darcs to revert change #23.
>> It also irritates me that Git insists that even unrelated changes must
>> have a linear time ordering.
>
> No they don't. Indeed, if unrelated changes had to have a linear time
> ordering, you wouldn't get merge conflicts at all, would you?
This doesn't make any sense to me at all...
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 21/04/2011 06:54 PM, Darren New wrote:
> OK. All the confusions you think you're seeing in git are due to Darcs
> just not being able to do what git does.
>
> When Darcs gets a merge conflict, it just doesn't apply *either* patch.
>
> "Darcs escapes this problem by ignoring those parts of the patches that
> conflict."
Interesting. And here I was thinking it marks the conflicting parts of
the files for you so you can go fix it...
> "If the conflict is with one of your not-yet-published patches, you may
> choose to amend that patch rather than creating a resolve patch."
>
> And that's exactly what the "git merge" command does.
I thought "git merge" just combines changes, not resolves conflicts.
> "This is how a project with many contributors, but every contribution is
> reviewed and manually applied by the project leader, can be run." This
> is the bit about sending email you were talking about. git can work that
> way, and the terrible "merge" problems you're talking about are handled
> the same way: the guy getting the patches fixes the merge.
>
> What Darcs apparently can't do is support any way of doing distributed
> development with an authoritative repository *without* someone dedicated
> to fixing the merge conflicts. That's where the whole rant you're
> talking about came from.
I wasn't even talking about conflicts. I'm talking about the fact that
if the central repo changes, even in a way which does *not* conflict
with your changes, you still have to update your local repo, remerge all
the changes, and try again.
> I'll grant you that Darcs is definitely simpler, but I think it's less
> capable also, and that's the primary place the simplicity comes from.
I disagree, but I don't think this argument is going anywhere productive
right now.
> The Darcs replace command is interesting, but I'm not sure how well that
> would work in practice, especially in languages with complex scoping.
Yeah, it's only really useful for global names (e.g., functions or
types). If you've got a dozen functions with a variable named "x1" and
you want to make it "x_in" in one of them... yeah, good luck. Really,
you're going to have to sort it out by hand.
What should *really* happen is that Darcs looks at your edits and
*detects* that it's a find-and-replace affecting only certain lines, and
record that. But anyway...
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/23/2011 3:47, Orchid XP v8 wrote:
> I will admit, a GUI is probably a superior interface for doing this. (And
> I'm not aware of anyone having built a GUI for Darcs.)
There's a couple of standard git guis that come with git. One for really
sophisticated exploration of the history and one for manipulation of the
repository and working directory.
>> Which means you can't (for example) stop in the middle when you realize
> True enough. It would be nice if you could make Darcs remember which changes
> you selected last time around. (Then again, some of those changes might no
> longer exist next time you run Darcs...)
Git also does some various funky things. Like if you're doing merges and you
get a conflict, it leaves extra information in the "index" part to keep
track of which merge conflicts you've fixed and which you haven't. The
"bisect" routine (which Darcs has as well) uses that area to track stuff. Etc.
Of course, in git they'll still exist, because it gets copied into the
repository when you say "remember this change", not when you say "commit".
> You still have to manually make Git combine all 200 changes into one
> version, and then commit that.
But if they don't conflict, then that's one step. (Well, one step for each
of the 200 changes, just like in Darcs.)
> It just seems like an undless cycle of merges
> trying to keep everything straight, generating an ever more tangled history
> behind it.
No, it doesn't really work like that. Like I said, of the Linux repo,
there's 240,000 commits, of which 15,000 are actually merges. Git has a lot
of ways of editing the history to make it look simpler, which is where all
the perceived complexity of git lies.
> It's more the conceptual annoyance of having to record every merge operation
> as a new version of the entire repo, even if you only changed one line. That
> seems really clumsy to me.
Just so you know: There's merges, and there's commits. A commit is the same
as what Darcs calls a changeset. A merge is when you have two different
branches that you're putting into one repository.
In git, a commit is a trivial thing. You say "git commit" just like in Darcs
you say "darcs record." I do dozens of git commits a day just farting
around with my own programs.
In git, a merge is when you take two branches and apply the changes from one
branch to the other branch. In Darcs, this would involve two separate
repositories, one being updated from the other. It's about the same
complexity as that.
>> Honestly, I'm not sure I see any advantage of Darcs over git.
>
> Likewise, but inverted.
Well, here's the advantages of git, so far:
Git can keep multiple branches in one repository, which makes it easier to
move changes around, look at history, try out new combinations of changes, etc.
Git keeps track of where changes came from and when, so if I pull in a
2-month-old changeset and it breaks something, I can figure out when I
pulled it in vs when you wrote the change. (I didn't see that in Darcs, but
I didn't look very close.)
git keeps a history that can tell me what changes rely on other changes
semantically, not just syntactically. (I.e., the thing Darcs tries to do
with the retoken command, except that really only works for one specific
type of semantic change and still only works syntactically.)
git doesn't have to spend tens of hours applying 240,000 change sets to the
repository in order to give me the latest version of the files.
git has all kinds of sweet tools to manipulate the repository in ways Darcs
can't very easily. (Basically, you'd have to tell Darcs to make a bunch of
snapshots, then fiddle with the snapshots the same way git would fiddle with
the repository itself.)
> Still, until GHC moves from Darcs to Git, I won't have to actually care, so
> I guess it doesn't really matter.
Not unless you wind up working on a project that uses git.
I think 90% of your confusion with git is you're trying to think about it
like you think about Darcs, in terms of a series of changes, which is why I
posted the original link in the first place.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/23/2011 3:25, Orchid XP v8 wrote:
> On 22/04/2011 12:39 AM, Darren New wrote:
>> On 4/21/2011 13:07, Orchid XP v8 wrote:
>>>> Well, git *does* store stuff in files, so technically you could copy the
>>>> files. But by "copy files" I mean "use git to copy the new files." As
>>>> in, "you don't have to run any diffs or patches or anything".
>>>
>>> Oh, I see.
>>
>> Actually, you can store a repository on an http server and give everyone
>> read access to the repository, and everyone can update from there. So,
>> yeah, it basically *can* be done by copying the right files from one
>> place to another.
>
> I've done this with Darcs. Unfortunately, to update the repo, you basically
> delete it off the server and copy the current version over there - which is
> tedious. (By default, the server would also have a working copy - which is
> pointless. Fortunately, if you delete that bit, it still works.)
Yeah. You can't update git via http, but you can update a repository stored
on the server and then there's a command called "update-http-hook" or
something that refreshes the index that clients use to quickly find files on
the http server. (I.e., the clients need to be able to get a directory
listing of all the heads of branches, which is something an http server
never really standardized.)
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
>> You still have to manually make Git combine all 200 changes into one
>> version, and then commit that.
>
> But if they don't conflict, then that's one step. (Well, one step for
> each of the 200 changes, just like in Darcs.)
Except that you have to record the order you did it in, even though it
doesn't actually matter.
>> It just seems like an undless cycle of merges
>> trying to keep everything straight, generating an ever more tangled
>> history behind it.
>
> No, it doesn't really work like that. Like I said, of the Linux repo,
> there's 240,000 commits, of which 15,000 are actually merges. Git has a
> lot of ways of editing the history to make it look simpler, which is
> where all the perceived complexity of git lies.
I might suggest that if you *need* the ability to edit history to make
it look simpler, you're doing it wrong.
> Just so you know: There's merges, and there's commits. A commit is the
> same as what Darcs calls a changeset. A merge is when you have two
> different branches that you're putting into one repository.
>
> In git, a commit is a trivial thing. You say "git commit" just like in
> Darcs you say "darcs record." I do dozens of git commits a day just
> farting around with my own programs.
>
> In git, a merge is when you take two branches and apply the changes from
> one branch to the other branch. In Darcs, this would involve two
> separate repositories, one being updated from the other. It's about the
> same complexity as that.
Except that with Darcs, you don't need to record the fact that a branch
ever even existed. You just record what was done with the actual file
contents.
> Well, here's the advantages of git, so far:
>
> Git can keep multiple branches in one repository
I don't know how that works, but if you mean you can have multiple
working copies, then yes, I guess that could be quite useful.
> Git keeps track of where changes came from and when, so if I pull in a
> 2-month-old changeset and it breaks something, I can figure out when I
> pulled it in vs when you wrote the change.
I'm not sure I see why you would need that information, but OK, Darcs
can't tell you about that.
> git keeps a history that can tell me what changes rely on other changes
> semantically, not just syntactically.
How?
> git doesn't have to spend tens of hours applying 240,000 change sets to
> the repository in order to give me the latest version of the files.
Neither does Darcs.
> git has all kinds of sweet tools to manipulate the repository in ways
> Darcs can't very easily.
Ways such as what?
>> Still, until GHC moves from Darcs to Git, I won't have to actually
>> care, so I guess it doesn't really matter.
>
> Not unless you wind up working on a project that uses git.
That isn't going to happen.
1. I will never wind up "working" on anything that's version-controlled.
I am apparently doomed to spend the rest of my /working/ life rebooting
people's PCs because Word crashed, rather than doing interesting coding
tasks.
2. If we're talking about hobby projects, obviously I'm going to pick
one that uses my preferred tools.
> I think 90% of your confusion with git is you're trying to think about
> it like you think about Darcs, in terms of a series of changes, which is
> why I posted the original link in the first place.
It's not that I don't understand the difference between Git and Darcs.
It's that I can't begin to comprehend how what Git does can work. It
just seems such an obviously stupid way to approach the problem.
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/23/2011 4:08, Orchid XP v8 wrote:
> Not to my knowledge, no. It records who created a patch and when, but not
> when it was applied to any particular repo.
That seems a problem to me, yes. :-)
>> So if, for example, I'm working, and everything's good, and I take some
>> patches from you, then work some more, then take some patches from Sam,
>> then work some more, then run my test and it fails, can I figure out
>> that it was Sam's patches, even if he created those patches before I
>> even cloned the repository in the first place?
>
> In that case you're presumably going to revert patches until the problem
> goes away. Maybe one patch broke something, maybe its an interaction of
> several patches. You turn patches on and off until you figure out what's up.
Sure. But I can't tell after the fact when I sucked Sam's patch in, so if
Sam wrote the patch 2 months ago and I only started seeing the problem a
week ago, it's not obvious that it might actually be Sam's patch.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/23/2011 4:15, Orchid XP v8 wrote:
> Hint: Look up "darcs fetch". It fetches changes without applying them.
OK. It sounded to me like Darcs had one working copy and one repository, and
applying all the patches from the repository would give you a "pristine"
version, and the working version was differenced from the pristine version
to get the next patch.
So it sounds like you can't actually fetch changes without having them
actually be relevant to the working directory. If you make the WD match the
repository, then fetch changes without applying them, then ask for diffs,
isn't Darcs going to tell you the WD doesn't match the repository? If your
change added the line ABC, isn't darcs diff going to tell you you removed
ABC from the WD?
So what do you mean "apply"? What happens if your WD matches the
repository, then you do a fetch, then you say "record", what does Darcs do?
> It's the fact that you have to commit the merged version of the file as a
> new version. Every time you merge in a new change, you have to commit a new
> version. It just seems clunky and unecessary.
It's mildly clunky, but it tells you which changes are applied to which
snapshots. git does this because it records files, not changes, yes. It's
probably the part of git that people like least.
What happens if I have a number of tags in the repository (v1 and v2 and the
current v3 under development), and I want to apply a bugfix to the v1
version that doesn't apply to the later versions? How do I do that?
>> In Darcs, that's one step it seems - I can't get changes
>> from you *without* applying them to the work I'm doing.
>
> Not true. You can get changes without applying them. It's just that usually,
> you want to fetch changes to, you know, *use* them.
But if you have multiple branches, like V1, V2, and V3beta1, and someone
sends you a new feature for V3beta1, you don't want to apply that to V1 or
V2. If you get a bugfix for V1 and make V1.01 you might or might not want
to apply that to V3.
> Then why fetch the changes at all? Why not wait until you're actually ready
> to apply them?
Welcome to DVCS!
Plus, what do you mean "apply"? git stores multiple branches in one
repository. I could as easily ask you "why wouldn't you apply every change
you get to every copy of the Darcs repository for your program?" The answer
is "because maybe I don't want to change that repository."
I might want to push changes up to a shared repository so everyone in the
team can work on them, but I don't want it going into production or even to
people not on the team.
I might want the development branch to apply the changes but not the branch
that's half way thru QA testing.
>> So if you have something like Linux, where there's a new release every
>> few months, you need a complete repository for every release.
>
> Um... why?
Because you don't want all the changes applied to old versions of Linux?
Maybe Darcs tags would do the trick there?
>> And if I
>> fix a bug in an old release and you want to incorporate that bug fix
>> into newer releases, what do you do?
>
> Oh, I see. You mean if you actually have multiple versions of something
> being developed concurrently? Yeah, in that case you'd have to move the
> changeset from one branch to another and hope it works.
And *that* is exactly what a git merge is. You don't do a merge every time
you record a change. You only do a merge when you're actually, you know,
merging two sets of changes into one branch of the repository.
A git branch is like a darcs repository.
A git commit is like a darcs record.
A git merge is like a darcs fetch-and-apply.
>>> You then have to
>>> *create a new commit object* representing this new combined state.
>>
>> Yes.
>
> So if two people have the same repo, and they merge in change X and then
> later merge in change Y, and then the other guy merges in change Y first and
> later change X, they now apparently have conflicting histories.
Well, you don't really "merge in" a change. You merge a change from one
branch to another. You're not merging changes, you're merging branches.
You're not going to have conflicting histories. Indeed, you *can't* have
conflicting histories because history is immutable. You might have different
histories in each repository, but that's like having two Darcs repositories
cloned from the same source but with a different set of changes in each.
So if we start at the same version, and you make three changes, and I make
three changes, and assuming there's no conflicts, when I merge your changes
into mine, I'll get the same files as if you merged my changes into yours,
so we'll both wind up with the same files, except yours will say yours
committed it and mine will say mine committed it.
Normally the only reason you'd both merge in the other person's changes is
if you're both going to keep developing on your own. Otherwise, what you
ought to do is I merge in your changes, and then give you the result to
continue from.
Obviously it's not a problem if you don't later combine the repositories
into one history. If you do, then you'll have two branches, and your pointer
will point to yours, and mine will point to mine, and they'll each be
pointing to a separate commit with the same files in each one (assuming
they're identical contents and that you didn't resolve merge conflicts
differently than I did).
> (And all
> because Git wants to pretend that everything happens in linear order.)
No, git very specifically does *not* pretend things happen in a linear
order. There's all kinds of tools to examine the DAG of dependencies
between versions.
> How do you get out of that?
If it comes time to turn your work and my work into the same branch, you
just merge the two branches in both repositories and work from there on,
which won't cause any merge conflicts because they're all the same files.
It's a DAG.
It would be similar in Darcs as if I set a tag in my repo and you set one in
your repo, then we merged them, and now you ask "how do you make a tag that
incorporates all the changes?" Well, if I took your changeset and you took
my change set, we'd have two tags. Make a third tag that points to the
combined set.
> As I say, it just annoys me that you have to assign an arbitrary ordering to
> changes.
It's not really arbitrary. Darcs just ignores the fact that they aren't.
If I add function AAA to file BBB and record that change, then add a call to
function AAA into file CCC, then record that change, would that really be
independent unordered changes in Darcs? That seems awfully fragile.
>> Say you have 25 changes in your Darcs repository, and you want a version
>> that applies every change except #23. What do you put in the repository
>> to represent that?
>
> You ask Darcs to revert change #23.
And how do you do that? "darcs revert" doesn't do that. "darcs unrecord"
throws the change away entirely. Obliterate deletes the patch entirely as well.
So it sounds like the actual answer is "you have to clone the entire
repository, then you have to obliterate the patch, and then you have to
rebuild the working directory from the new repository, and *then* you're
done, yes?
>>> It also irritates me that Git insists that even unrelated changes must
>>> have a linear time ordering.
>>
>> No they don't. Indeed, if unrelated changes had to have a linear time
>> ordering, you wouldn't get merge conflicts at all, would you?
>
> This doesn't make any sense to me at all...
Patches within any one branch have a linear time ordering. But git handles
multiple branches in the same repository. So changes between branches aren't
linearly ordered.
If changes were all linearly ordered, you'd never have to resolve merges,
because you could never have two independent patches trying to be applied to
the same place in the file, because one would definitely come before the
other, right?
Even on the same branch, while the changes are recorded linearly, if there
aren't conflicts between them, it's trivial to rearrange the order however
you want. Just like in Darcs you can't rearrange a patch that deletes a line
to be applied earlier than the patch that creates the line, so there's a
partial ordering in Darcs also. Same in git - it's a DAG with a partial
ordering. It's not linear.
For example, attached is a picture of the last few dozen commits to the
Linux kernel.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
Attachments:
Download 'image1.png' (172 KB)
Preview of image 'image1.png'

|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
On 4/23/2011 4:14, Orchid XP v8 wrote:
> Seems more like because it's so difficult to merge two files,
It's no more difficult in git than in darcs. It's exactly the same process.
You apply one set of diffs, then the other. If there are no conflicts,
you're done. If there are conflicts, you fix them, and you're done.
If there's a branch called "newstuff" and I'm working on "master", and I
want to merge in the changes from newstuff, I say
git merge newstuff
git commit
If file ABC has a conflict in it, I say
git merge newstuff
vi ABC
git add ABC
git commit
> after you've
> done it you have to save it to prevent you having to redo all that complex
> hard work.
Yes. Except it's exactly as hard and complex as in Darcs.
> And in the process, all change application is forced to become
> strictly linear.
git doesn't record changes, so no, change application isn't forced to become
linear.
> Apparently this doesn't stop people working on the Linux kernel. But it
> seems really clumsy to me.
It seems clumsy because you keep thinking git is recording changes. You keep
thinking of "changes" instead of "versions".
>> I'm not sure I'd want to check out something from Darcs that has a
>> quarter million patches in it and wait for Darcs to apply them all one
>> by one. How well does it handle that?
>
> You're aware that Darcs keeps a cached copy of the latest state of all the
> files, so it doesn't have to recompute them, right?
Sure.
> Last time I tried downloading the repos for GHC, it was dominated by network
> latency. Processor usage was almost non-existent. It just takes a long time
> to shift gigabytes of data over a slow ADSL link. Just as it would if I had
> downloaded a Zip file of the source code with no history data at all.
Yep. But you only got the latest version. If you wanted to get every tag,
you wind up copying all those files again anyway.
--
Darren New, San Diego CA, USA (PST)
"Coding without comments is like
driving without turn signals."
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|  |
|
 |
>> In that case you're presumably going to revert patches until the problem
>> goes away. Maybe one patch broke something, maybe its an interaction of
>> several patches. You turn patches on and off until you figure out
>> what's up.
>
> Sure. But I can't tell after the fact when I sucked Sam's patch in, so
> if Sam wrote the patch 2 months ago and I only started seeing the
> problem a week ago, it's not obvious that it might actually be Sam's patch.
Or, to summarise, "if I let my repo get 6 months out of date with the
upstream and then pull everything in at once and try to figure out why
it broke, it'll be quite difficult". My general reaction being "don't do
that", but OK, I guess it's a valid complaint...
--
http://blog.orphi.me.uk/
http://www.zazzle.com/MathematicalOrchid*
Post a reply to this message
|
 |
|  |
|  |
|
 |
|
 |
|  |
|
 |